Qwen3.5-397B-A17B-VQ-2.6bpw · Model Card
Qwen3.5-397B-A17B-VQ-2.6bpw: Model Card
Written by Noah Zelezny, published under apache-2.0, revision 5ff492af70af, read 2026-10-02. Shown as written; SAVRN's own facts about this model are on its page.
105.5 GiB text weights — the balanced build. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 111.8 GiB.
It lands in the same size class
as the strongest community quant at this rate; the table below is the
comparison. A vector-quantized build of
Qwen3.5-397B-A17B for Apple
Silicon. Stock mlx-lm, no patches — the VQ runtime ships inside the
checkpoint as model.py. It sits between the VQ-2.4bpw daily driver and the
VQ-3bpw quality build, and needs the same hardware class as the latter.
Changelog
2026-10-01 — vq-skipzero: dead rows dropped
12.43% of the 64-element weight groups in the Qwen3.5-397B teacher sit in output rows whose weights are about 1e-29. The vq-skipzero format drops those rows' codes and scales on disk and in memory; they output exact zeros, and the live rows are byte-identical to the previous revision (vqlab sz-check, every module). Text weights go from 119.21 to 105.54 GiB (-13.67 GiB), and resident memory drops by about the same amount. Paired 3-corpus KL against the previous revision (12288 tokens each, 2026-09-29): no corpus differs at |t|>2.
2026-10-01 — runtime refresh; Knurlogic
The bundled model.py is updated to the current VQ runtime; weights are
unchanged. The run, multi-Mac, drafting and vision instructions now use
Knurlogic.
2026-09-20 — correction: layers 57-59 were affine, and now are not
What was wrong. This rung was described as a uniform VQ build, but the expert modules of layers 57, 58 and 59 — 9 of 180 — were never fitted and fell through to the base affine 3-bit, group 64. The fit had been invoked over layers 0-56 of a 60-layer model, and the off-by-three then traveled as a constant into every downstream tool. Those three layers therefore carried more bits than the rung's nominal rate, not fewer, so the published quality numbers were mildly optimistic.
What changed. The 9 modules are refitted at this rung's own geometry from the bf16 teacher, so the build is now genuinely uniform across all 60 layers. It is slightly smaller and slightly worse, and both are expected: the correction removes bits from late layers, where they do the most good.
The KL table below has been re-measured on the corrected weights — same instrument, same teacher cache, same positions as before. Sizes on this card are now stated consistently as text weights (the number that governs whether it runs); the vision tower (+0.85 GiB) and the optional MTP draft head (+5.4 GiB) are real extra bytes and are called out wherever a size appears. Earlier revisions of this card mixed those three totals in one column.
This is a correction to the numbers, not to how the model is used: the runtime, the config surface, the vision tower and the MTP head are all unchanged, and text and image generation were both re-verified on a 2-node exo pipeline after the rebuild.
2026-09-17 — the full Qwen3.5-397B-A17B VQ ladder on one KL instrument
KL to this family's own bf16 teacher (cached top-64), 12288 tokens, on all three house corpora — prose, public code and literary — every rung scored on the same cache and the same positions, through a streamed pass validated against a direct full-model forward to all printed decimals.
| build | GiB (text) | prose | code | literary | mean | top-1 |
|---|---|---|---|---|---|---|
| VQ-2.2bpw (v2, mixed) | 100.0 | 276.4 | 98.6 | 183.1 | 186.0 | 84.8% |
| VQ-2.4bpw | 95.7 | 230.4 | 87.7 | 134.1 | 150.7 | 86.2% |
| VQ-2.6bpw (this) | 105.5 | 164.1 | 58.1 | 60.5 | 94.2 | 88.5% |
| spicyneuron 2.6bit (affine) | 120.6 | 332.1 | 98.4 | 241.0 | 223.8 | 83.5% |
| VQ-3.1bpw | 125.3 | 91.4 | 32.3 | 16.3 | 46.7 | 91.7% |
Millinats per token; lower is closer to bf16. The KL and top-1 rows were scored on the full-row revisions; sizes are the current vq-skipzero ones. These rows are VQ rungs only and are not comparable to the affine-comparator table further down, which was measured prose-only at 2048 tokens on an earlier instrument.
Rank by KL, not perplexity. On this family perplexity is not monotone in quality: rungs that are measurably closer to the teacher can read a higher perplexity, because a damaged model can score better than bf16 on a finite sample. KL on cached teacher logits does not have that failure mode.
2026-09-09 — runtime refresh
Runtime refresh. The bundled model.py is updated so downloaders run
exactly the code that was benchmarked; dense bundles now carry both runtimes.
- Faster prefill on affected geometries, measured per rung: a device-codebook kernel arm for large-codebook geometries (up to 1.46x on affected rungs), a ragged-subvector relaxation (up to 1.34x on affected rungs), a fused d8 arm, and a routing memo (≈2%). No blanket speedup is claimed across the lineup — gains apply only where the geometry engages the new paths.
- Output quality is unchanged: the kernel changes are bit-identical or 1-ULP-equivalent, and the routing memo is bit-identical (logits checksum verified).
- Speculative decoding (MTP): on repos that ship
mtp-head-q6.safetensors, the sidecar works with the exo fork branchmtp-stage1(github.com/noahzelezny/exo) — launch each node withexo --mtp(or setEXO_MTP=1). Note: with MTP enabled, exo serves requests sequentially (the batch engine has no MTP path), so leave it off for concurrent / multi-agent workloads.
2026-09 — bundle refresh (runtime update landed)
Bundle refresh (runtime update landed).
model.py updated: refreshed
VQ expert kernels (verified equivalent), a default buffer-cache ceiling
(VQLAB_CACHE_LIMIT_GB overrides, 0 disables) so transient prefill
allocations return to the OS as they free — long-prompt peak stays near
resident instead of a multiple of it — and dual-runtime loading. Weights
unchanged; prior revision pinnable. This rung exceeds the single-box gate
bar, so the update shipped through the release gate's cluster smoke: a
real generation on a 2-node exo pipeline, with the peer rank's copy
identity-checked before the generation counted.
Measured results
See also the 2026-09-17 entry in the changelog above, which places this rung against the rest of its family on the three-corpus KL instrument. The table below is the earlier prose-only measurement and is not comparable to it row-for-row.
Scored on this exact artifact with an unmodified mlx-lm install, on the same
two corpora and the same harness as every comparator below:
| this model (105.5 GiB text) | spicyneuron 2.6bit (120.6 GiB text) | |
|---|---|---|
| wikitext perplexity (raw, prefix-8192)* | 2.5634 | 3.1843 |
| code perplexity (mixed-language)* | 2.6123 | 2.6667 |
19.5% better on prose and 2.0% better on code, at 15.1 GiB less on disk.
*wikitext/code perplexity here is the older external-corpus instrument, measured pre-correction; the KL table above is the current ranking, re-measured on the corrected weights.
Both margins are large relative to the fit-to-fit noise we can measure: 24x and 3.1x respectively. One caveat on those multiples, because it is the kind of thing that is easy to leave out — the noise floor they are quoted against was measured at a different codebook size (256, not 512). No floor has been measured at this geometry. Floors in this project have widened every time they were measured more carefully, so treat 24x as "comfortably real" rather than as a precise figure. The prose gap is not in doubt at any plausible floor; the code gap is the smaller of the two and would be the first to become uninteresting if this geometry's floor turned out to be wide.
The size comparison, measured. This artifact carries the full vision tower at bf16 (333 tensors, 0.849 GiB); the spicyneuron builds are text-only and carry none. Comparing text weights to text weights: this build is 105.5 GiB against their 120.6 — 15.1 GiB smaller, and ahead on all three corpora of the KL table above. Before the 2026-09-20 geometry correction it was 121.4 against their 120.6 (+0.8 GiB), so the correction moved this rung from slightly larger to slightly smaller than the comparator; the 2026-10-01 vq-skipzero revision removed another 13.7 GiB.
Sizes on this card are model weights, including the vision tower. The
repo also ships an optional mtp-head-q6.safetensors drafting sidecar
(5.41 GiB) which mlx-lm does not load; a full snapshot_download fetches
it, so budget for it separately.
Task benchmarks
Not yet measured on this artifact. Two of the siblings above carry HellaSwag/PIQA/WinoGrande numbers; this build has not been run through that harness, and reporting a sibling's task scores here would be exactly the substitution this project refuses to make. They will be added once the suite has been run under the same harness (lm-eval 0.4.12, layer-streaming loglikelihood scorer, 0-shot, first 1000 items per task).
Hardware
This build has not been verified on a single 128 GB machine. Use either:
- a single Apple Silicon machine with ≥ 192 GB unified memory, or
- two or more Macs with Knurlogic (see Run it). This rung was verified on a 96 GB + 128 GB exo cluster over Thunderbolt.
Verified serving on a 2-node ring: placed and serving in 99 seconds, 800-token coherent generation, and three graded known-answer probes returned correct. No throughput figure is published here — we measured placement and correctness, not tokens per second, and an unmeasured number is worse than none.
This artifact has not been verified single-node. Any re-verification has to be a 2-node ring.
Speed
Measured 2026-10-01 on this revision.
Split across two Macs with Knurlogic (M3 Ultra 96 GB + M4 Max 128 GB, pipeline over Thunderbolt), against Qwen3.5-397B-A17B-MLX-2.6bit on the same two Macs: 3 pairs, a prompt of about 1,900 tokens, 128 generated tokens, no drafting.
| this model vs Qwen3.5-397B-A17B-MLX-2.6bit | |
|---|---|
| decode | 0.83× |
| prefill | 0.88× |
A two-Mac split's absolute speed depends on the link, so only the ratio is quoted.
The ratios above were measured on the previous full-row revision. Generation on this revision was verified 2026-10-01 on the same two Macs at 38.2 tok/s decode (short prompt, 8-40 tokens; not a paired benchmark).
Run it
With Knurlogic, which works out the settings and checks the model fits before loading it:
pip install knurlogic
hf download TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.6bpw --local-dir ~/Knurlogic/Models/Qwen3.5-397B-A17B-VQ-2.6bpw
knurlogic serve ~/Knurlogic/Models/Qwen3.5-397B-A17B-VQ-2.6bpw
Chat at http://127.0.0.1:8080/, or point any OpenAI-, Anthropic- or
Ollama-compatible client at it (model local). knurlogic ui opens a page
with a Launch button for every model on disk instead.
At 120 GB on disk this needs a Mac with more memory than that, or two Macs:
install the same Knurlogic version and the model on each, run
knurlogic ui --host cluster on both, and launch it from the page with both
machines selected. Knurlogic splits the layers between them.
Images work through Knurlogic directly; mlx-vlm is not needed.
The VQ runtime ships inside this repo as model.py (declared by
model_file in config.json), and Knurlogic runs the file the model ships.
Without Knurlogic:
pip install 'mlx-lm>=0.31.3'
python -m mlx_lm generate \
--model TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.6bpw \
--prompt "Explain vector quantization briefly." \
--max-tokens 1000
Speculative decoding (MTP)
mtp-head-q6.safetensors (5.4 GiB) is the model's own MTP draft head —
the same file validated on the 2.2bpw rung (acceptance 0.72, single-box
via vqlab serve --sidecar). Cluster speculative decoding is live: Knurlogic loads the shipped head
and drafts automatically, on one Mac or split across several, with nothing
to enable (outputs exactly the base model's via rejection sampling).
Measured on this family's 2.6bpw rung on a 2-node exo pipeline: acceptance 0.85, throughput at parity
with exo's stock decode — the head drafts well, but this family's stock
pipeline decode does not degrade with generation length, so there is
little for speculation to recover. Expect
parity.
Vision
The artifact includes the full 333-tensor vision tower at source precision
(0.85 GiB). mlx-lm is text-only for this architecture and ignores it;
Knurlogic loads it from this folder directly, without mlx-vlm.
The sizes quoted above are the download: they include this tower. Because
mlx-lm does not load it, resident memory runs ≈0.85 GiB below the disk
figure.
Siblings
All from the same skeleton and recipe, all scored the same way:
| size | wikitext | code | needs | |
|---|---|---|---|---|
VQ-2.2bpw |
100.0 GiB | 3.0591 | 2.6728 | 128 GB Mac, roomy |
VQ-2.4bpw |
95.7 GiB | 2.7655 | 2.6383 | 128 GB Mac, tight |
VQ-2.6bpw (this build) |
105.5 GiB | 2.5634 | 2.6123 | ≥192 GB or cluster |
VQ-3.1bpw |
125.3 GiB | 2.3410 | 2.5963 | ≥192 GB or cluster |
If you can run this build you can run VQ-3bpw, which is better on both
corpora for 19.8 GiB more. Take this one if those gigabytes are worth more to
you than the quality difference — on a 192 GB machine it leaves roughly 84 GiB
free against the 3bpw build's 64.
Methodology
Mixed precision by layer sensitivity. Attention, MoE routers, embeddings and the output head stay at higher precision — a small fraction of the parameters, but errors there propagate through every token. The MoE experts are ≈85% of the model and individually far more tolerant, so they absorb the aggressive quantization.
Vector quantization instead of scalar rounding — the part that is different. Scalar 2-bit gives each weight 4 rigid levels; over a group of 4 weights that is 256 fixed grid combinations. This build learns a codebook of joint 4-weight patterns and stores one index per group. Each 4-weight subvector stores one 9-bit index into a per-tensor 512-entry fp16 codebook, with an fp16 scale per (row, 64 weights) — 2.5 bits/weight stored in the expert region. At the same bits the codebook's entries sit where the weight distribution actually is, rather than on a uniform lattice.
Every expert tensor uses this one geometry; there is no mixed allocation and no per-layer schedule. Flat rungs are the reference points in this lineup because no mixed-allocation build we measured beat the flat rung at its own size.
Codebooks are fit in pure weight space — k-means over the weight subvectors, no Hessian, no activation statistics, no calibration corpus. Calibration-fitted methods we tested (GPTQ- and DWQ-style) reduced layer error while making end-to-end perplexity worse on this architecture, and they bias the result toward whatever text the calibration set contains.
The fit is not seeded. k-means draws an unseeded subsample, so this artifact is reproducible in recipe and geometry but not bit-for-bit. That is why margins here are quoted against a measured fit-to-fit floor rather than against a repeated build.
Sub-byte bit-packing. Codes are packed into uint32 words (row-local, 32-code blocks) rather than padded to whole bytes, which is what makes the non-byte-aligned size possible. Packing is a pure representation change.
How it was evaluated. Perplexity on two corpora — raw wikitext
(prefix-8192) and a mixed-language code corpus — scored with an unmodified
mlx-lm on the same harness used for every comparator here. Two corpora
because this family shows real domain asymmetry: larger codebooks buy far
more on prose than on code, so a single-corpus number would misrepresent the
trade.
Verification
Release gates passed on this artifact before upload: file, index and tokenizer checks, a verbatim match between the bundled runtime and its source, and a generation smoke through the shipping runtime on Apple Silicon. The upload path runs the gate itself and refuses to publish without it.
Limitations
- No throughput measurement — see Hardware.
- Perplexities are corpus-specific. Compare only against models scored on the same files, never across harnesses.
- This is a thinking model: it spends tokens reasoning before answering.
Budget
max_tokensaccordingly.
Acknowledgment
spicyneuron's 397B quants are what made this model runnable on my hardware in the first place, and they were the reference this work was measured against throughout. This release is offered in that spirit: the method, the failures as well as the wins, and comparator numbers re-measured on one harness so the claims can be checked rather than taken on trust.
Paper
The method, the full model ladder, the negative results, and the measurement rules behind every number here: Data-Free Vector Quantization Beats Affine Quantization at Matched Bytes Below 6 Bits (CC BY 4.0) · code: VQLab · web version: Space
Support this work
VQLab and these artifacts are built and released independently — the fits, the measurement harness and the published rungs are one person's compute and time. If they are useful to you:
- Sponsor: github.com/sponsors/noahzelezny
- Contract work: available for quantization and on-device inference work on Apple silicon — custom rungs, per-layer allocation for your model, or getting a checkpoint to run well on a Mac. Contact: [email protected]
Provenance
Base model: Qwen/Qwen3.5-397B-A17B — Apache-2.0. This is a quantized derivative and inherits that license; using it means accepting the base model's terms. Quantization: TheDrainFlorist, 2026.
Built with MLX and VQLab.