Qwen3.8-Flash-Next in the light DS4-IQ2 packaging of an IQ2 main (IQ2XXS gate/up, Q2K down padded to 768, embedded MTP block) plus an external demand-paged PLE Q41 sidecar instead of a ~95 GiB resident BF16 n-gram. Runs resident and zero-swap on a 64 GiB Apple Silicon box (measured on M5 Pro), 8K→220K context. - MTP (--mtp) adds ~+17% single-stream over MTP-off; draft acceptance ~67%. - Adaptive draft depth (engine env DS4QWEN4MTPDEPTH, default auto): drafts a 2nd token on deterministic/structured/code continuations for a further +5–6%, and falls back on free-form prose so it never regresses. Output stays autoregressive-exact (verify only commits argmax-matching drafts → zero quality…
Independent publisher
DongNH
dongnhdev
Models in Library1
Datasets in Library0
Models on Hugging Face1
Followers2