Qwen3.5-397B-A17B-VQ-2.4bpw · Model Card
Qwen3.5-397B-A17B-VQ-2.4bpw: Model Card
Written by Noah Zelezny, published under apache-2.0, revision 34136c9eaf83, read 2026-10-02. Shown as written; SAVRN's own facts about this model are on its page.
95.7 GiB text weights — the daily driver, runs on a single 128 GB Mac. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 101.9 GiB.
A vector-quantized build of Qwen3.5-397B-A17B
that fits and generates on one 128 GB Apple Silicon machine — no cluster,
no patches, stock mlx-lm.
Changelog
2026-10-01 — vq-skipzero: dead rows dropped
12.43% of the 64-element weight groups in the Qwen3.5-397B teacher sit in output rows whose weights are about 1e-29. The vq-skipzero format drops those rows' codes and scales on disk and in memory; they output exact zeros, and the live rows are byte-identical to the previous revision (vqlab sz-check, every module). Text weights go from 107.96 to 95.66 GiB (-12.30 GiB), and resident memory drops by about the same amount (measured peak 119.87 -> 107.06 GB). Paired 3-corpus KL against the previous revision (12288 tokens each, 2026-10-01): prose +0.16 mnats (t 0.26), code +0.89 (t 1.45), literary +0.43 (t 0.52); no corpus differs at |t|>2.
2026-10-01 — runtime refresh; Knurlogic
The bundled model.py is updated to the current VQ runtime; weights are
unchanged. The run, multi-Mac, drafting and vision instructions now use
Knurlogic.
2026-09-20 — correction: layers 57-59 were affine, and now are not
What was wrong. This rung was described as a uniform VQ build, but the expert modules of layers 57, 58 and 59 — 9 of 180 — were never fitted and fell through to the base affine 3-bit, group 64. The fit had been invoked over layers 0-56 of a 60-layer model, and the off-by-three then traveled as a constant into every downstream tool. Those three layers therefore carried more bits than the rung's nominal rate, not fewer, so the published quality numbers were mildly optimistic.
What changed. The 9 modules are refitted at this rung's own geometry from the bf16 teacher, so the build is now genuinely uniform across all 60 layers. It is slightly smaller and slightly worse, and both are expected: the correction removes bits from late layers, where they do the most good.
The KL table below has been re-measured on the corrected weights — same instrument, same teacher cache, same positions as before. Sizes on this card are now stated consistently as text weights (the number that governs whether it runs); the vision tower (+0.85 GiB) and the optional MTP draft head (+5.4 GiB) are real extra bytes and are called out wherever a size appears. Earlier revisions of this card mixed those three totals in one column.
This is a correction to the numbers, not to how the model is used: the runtime, the config surface, the vision tower and the MTP head are all unchanged, and text and image generation were both re-verified on a 2-node exo pipeline after the rebuild.
2026-09-17 — the full Qwen3.5-397B-A17B VQ ladder on one KL instrument
KL to this family's own bf16 teacher (cached top-64), 12288 tokens, on all three house corpora — prose, public code and literary — every rung scored on the same cache and the same positions, through a streamed pass validated against a direct full-model forward to all printed decimals.
| build | GiB (text) | prose | code | literary | mean | top-1 |
|---|---|---|---|---|---|---|
| VQ-2.2bpw (v2, mixed) | 100.0 | 276.4 | 98.6 | 183.1 | 186.0 | 84.8% |
| VQ-2.4bpw (this) | 95.7 | 230.4 | 87.7 | 134.1 | 150.7 | 86.2% |
| VQ-2.6bpw | 105.5 | 164.1 | 58.1 | 60.5 | 94.2 | 88.5% |
| spicyneuron 2.6bit (affine) | 120.6 | 332.1 | 98.4 | 241.0 | 223.8 | 83.5% |
| VQ-3.1bpw | 125.3 | 91.4 | 32.3 | 16.3 | 46.7 | 91.7% |
Millinats per token; lower is closer to bf16. The KL and top-1 rows were scored on the full-row revisions; sizes are the current vq-skipzero ones. Sizes are text weights, so every row is like-for-like — spicyneuron is text-only and carries no vision tower. The affine 2.6bit build is larger than our 2.6bpw yet 2× worse on prose and literary, and this rung (2.4bpw) beats it on all three corpora at 24.9 GiB less. KL and top-1 agree: we lead on both.
Rank by KL, not perplexity. On this family perplexity is not monotone in quality: rungs that are measurably closer to the teacher can read a higher perplexity, because a damaged model can score better than bf16 on a finite sample. KL on cached teacher logits does not have that failure mode.
2026-09-09 — runtime refresh
Runtime refresh. The bundled model.py is updated so downloaders run
exactly the code that was benchmarked; dense bundles now carry both runtimes.
- Faster prefill on affected geometries, measured per rung: a device-codebook kernel arm for large-codebook geometries (up to 1.46x on affected rungs), a ragged-subvector relaxation (up to 1.34x on affected rungs), a fused d8 arm, and a routing memo (≈2%). No blanket speedup is claimed across the lineup — gains apply only where the geometry engages the new paths.
- Output quality is unchanged: the kernel changes are bit-identical or 1-ULP-equivalent, and the routing memo is bit-identical (logits checksum verified).
- Speculative decoding (MTP): on repos that ship
mtp-head-q6.safetensors, the sidecar works with the exo fork branchmtp-stage1(github.com/noahzelezny/exo) — launch each node withexo --mtp(or setEXO_MTP=1). Note: with MTP enabled, exo serves requests sequentially (the batch engine has no MTP path), so leave it off for concurrent / multi-agent workloads.
2026-09 — bundle refresh (runtime update landed)
Bundle refresh (runtime update landed).
model.py updated: refreshed
VQ expert kernels (verified equivalent), a default buffer-cache ceiling
(VQLAB_CACHE_LIMIT_GB overrides, 0 disables) so transient prefill
allocations return to the OS as they free — long-prompt peak stays near
resident instead of a multiple of it — and dual-runtime loading. Weights
unchanged; prior revision pinnable. This rung exceeds the single-box gate
bar, so the update shipped through the release gate's cluster smoke: a
real generation on a 2-node exo pipeline, with the peer rank's copy
identity-checked before the generation counted.
Measured results
The head-to-head against the affine comparator is the KL table in the changelog above — one instrument, text-weight sizes, spicyneuron scored on the same cache and positions. This rung is 24.9 GiB smaller than spicyneuron 2.6bit and ahead on all three corpora; it is the pound-for-pound build of the family.
Runtime, single M4 Max 128 GB (macOS, stock mlx-lm):
| load time | ≈60 s |
| resident memory | ≈95.7 GiB text weights (mlx-lm skips the tower); peak 117.7 GiB at 30k context, measured pre-correction on the full-row revision (a separate peak measurement went from 119.87 GB to 107.06 GB with vq-skipzero) |
| context verified | 30,031 tokens, zero swap growth |
| decode | ≈19–22 tok/s, flat from 512 → 14k context |
| prefill | ≈40–50 tok/s (chunked, as mlx-lm does natively) |
Perplexities are corpus-specific: never compare them across different corpora or eval harnesses, only against other models scored on the same files. The wikitext margin (13.2%) is much larger than the code margin (1.07%) — that asymmetry is real, so judge by your workload.
Speed
Measured 2026-10-01 on this revision.
Split across two Macs with Knurlogic (M3 Ultra 96 GB + M4 Max 128 GB, pipeline over Thunderbolt), against Qwen3.5-397B-A17B-MLX-2.6bit on the same two Macs: 3 pairs, a prompt of about 1,900 tokens, 128 generated tokens, no drafting.
| this model vs Qwen3.5-397B-A17B-MLX-2.6bit | |
|---|---|
| decode | 0.89× |
| prefill | 0.93× |
A two-Mac split's absolute speed depends on the link, so only the ratio is quoted.
The ratios above were measured on the previous full-row revision. Generation on this revision was verified 2026-10-01 on the same two Macs at 39.4 tok/s decode (short prompt, 8-40 tokens; not a paired benchmark).
Run it
With Knurlogic, which works out the settings and checks the model fits before loading it:
pip install knurlogic
hf download TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.4bpw --local-dir ~/Knurlogic/Models/Qwen3.5-397B-A17B-VQ-2.4bpw
knurlogic serve ~/Knurlogic/Models/Qwen3.5-397B-A17B-VQ-2.4bpw
Chat at http://127.0.0.1:8080/, or point any OpenAI-, Anthropic- or
Ollama-compatible client at it (model local). knurlogic ui opens a page
with a Launch button for every model on disk instead.
At 110 GB on disk this needs a Mac with more memory than that, or two Macs:
install the same Knurlogic version and the model on each, run
knurlogic ui --host cluster on both, and launch it from the page with both
machines selected. Knurlogic splits the layers between them.
Images work through Knurlogic directly; mlx-vlm is not needed.
The VQ runtime ships inside this repo as model.py (declared by
model_file in config.json), and Knurlogic runs the file the model ships.
Without Knurlogic:
pip install 'mlx-lm>=0.31.3'
python -m mlx_lm generate \
--model TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.4bpw \
--prompt "Explain vector quantization briefly." \
--max-tokens 1000
Speculative decoding (MTP)
mtp-head-q6.safetensors (5.4 GiB) is the model's own MTP draft head —
the same file validated on the 2.2bpw rung (acceptance 0.72, single-box
via vqlab serve --sidecar). Cluster speculative decoding is live: Knurlogic loads the shipped head
and drafts automatically, on one Mac or split across several, with nothing
to enable (outputs exactly the base model's via rejection sampling).
Measured on this family's 2.6bpw rung on a 2-node exo pipeline: acceptance 0.85, throughput at parity
with exo's stock decode — the head drafts well, but this family's stock
pipeline decode does not degrade with generation length, so there is
little for speculation to recover. Expect
parity; per-rung numbers for this artifact are not yet measured.
Vision
The artifact includes the full 333-tensor vision tower at source precision
(0.85 GiB). mlx-lm is text-only for this architecture and ignores it;
Knurlogic loads it from this folder directly, without mlx-vlm.
Siblings
This is the middle of a three-size family, all from the same skeleton and recipe, all measured the same way:
| size | wikitext | code | needs | |
|---|---|---|---|---|
VQ-2.2bpw (accessibility) |
100.0 GiB | 3.1706 | 2.6988 | 128 GB Mac, roomy |
VQ-2.4bpw (this build) |
95.7 GiB | 2.7655 | 2.6383 | 128 GB Mac, tight |
VQ-3.1bpw (quality) |
125.3 GiB | 2.3519 | 2.5987 | ≥192 GB or cluster |
(wikitext/code perplexity columns are the older external-corpus instrument, measured pre-correction; use the KL table above for the current ranking.)
Methodology
Mixed precision by layer sensitivity. Not all weights deserve the same bits. Attention, MoE routers, embeddings, and the output head stay at higher precision — they are a small fraction of the parameters but errors there propagate through every token. The MoE experts are ≈85% of the model and individually far more tolerant, so they absorb the aggressive quantization. A tail of later layers is also promoted above the expert baseline; measured layer-wise error showed depth matters, and the last layers repay the bits.
Vector quantization instead of scalar rounding — the part that is different. Scalar 2-bit gives each weight 4 rigid levels; over a group of 4 weights that is 256 fixed grid combinations. This build instead learns a codebook of joint 4-weight patterns and stores one index per group. Each 4-weight subvector stores one 8-bit index into a per-tensor 256-entry fp16 codebook. At the same bits, the codebook's entries sit where the weight distribution actually is, rather than on a uniform lattice — which is why this beats scalar quantization at matched size rather than merely matching it. Per-tensor codebooks, with an fp16 scale per (row, 64 weights), for 2.25 bits/weight stored in the expert region.
Codebooks are fit in pure weight space — k-means over the weight subvectors, no Hessian, no activation statistics, no calibration corpus. That is a deliberate choice: calibration-fitted methods we tested (GPTQ- and DWQ-style) reduced layer error while making end-to-end perplexity worse on this architecture, and they bias the result toward whatever text the calibration set contains. Weight-space fitting has no such domain preference.
Sub-byte bit-packing. Codes are packed into uint32 words (row-local, 32-code blocks) rather than padded to whole bytes, which is what makes the non-byte-aligned sizes possible at all. Packing is a pure representation change: the packed artifact's perplexities match its unpacked twin to four decimals of total negative log-likelihood on both corpora.
How it was evaluated. Perplexity on two corpora — raw wikitext
(prefix-8192) and a mixed-language code corpus — every number reproduced
bit-identically twice, scored with an unmodified mlx-lm. Two corpora
because this family shows real domain asymmetry: larger codebooks buy far
more on prose than on code, so a single-corpus number would misrepresent the
trade. Task-suite evals (HellaSwag/PIQA/WinoGrande and friends) have not
been run; only what is reported above is measured.
Verification
Release gates passed on this artifact before upload: file, index and tokenizer checks, a verbatim match between the bundled runtime and its source, and a generation smoke through the shipping runtime on Apple Silicon. The upload path runs the gate itself and refuses to publish without it.
Limitations
- Tight on 128 GB. The measured peak dropped from 119.87 GB to 107.06 GB with vq-skipzero, against ≈120 GiB usable. It runs; it is not roomy.
- This is a thinking model (Qwen3.5 family): by default it spends tokens
reasoning before answering. Budget
max_tokensaccordingly. - To split this model across two or more Macs, use Knurlogic (see Run it). Single-box users are unaffected.
Paper
The method, the full model ladder, the negative results, and the measurement rules behind every number here: Data-Free Vector Quantization Beats Affine Quantization at Matched Bytes Below 6 Bits (CC BY 4.0) · code: VQLab · web version: Space
Support this work
VQLab and these artifacts are built and released independently — the fits, the measurement harness and the published rungs are one person's compute and time. If they are useful to you:
- Sponsor: github.com/sponsors/noahzelezny
- Contract work: available for quantization and on-device inference work on Apple silicon — custom rungs, per-layer allocation for your model, or getting a checkpoint to run well on a Mac. Contact: [email protected]
Provenance
Base model: Qwen/Qwen3.5-397B-A17B — Apache-2.0. This is a quantized derivative and inherits that license; using it means accepting the base model's terms. Quantization: TheDrainFlorist, 2026.
Built with MLX and VQLab.