Qwen3.5-397B-A17B-VQ-2.2bpw · Model Card
Qwen3.5-397B-A17B-VQ-2.2bpw: Model Card
Written by Noah Zelezny, published under apache-2.0, revision f10167305589, read 2026-10-02. Shown as written; SAVRN's own facts about this model are on its page.
88.7 GiB text weights — the accessibility build, the roomiest fit on a 128 GB Mac. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 95.0 GiB. (v2, mixed geometry.)
v2 — updated 2026-08-22. This repository now serves a rebuilt artifact at the same size and the same bits per weight, with a different codebook geometry that measures better on both perplexity corpora. v1's numbers are kept below rather than quietly overwritten, and v1's bytes remain downloadable by pinning the previous revision:
snapshot_download("TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.2bpw",
revision="4554635165011f67e8166fd94d4bcc8cbf91401c") # v1
A vector-quantized build of Qwen3.5-397B-A17B
built to answer one question: how small can a 397B get and still be worth
running? 88.7 GiB text weights — it runs on a single 128 GB Apple Silicon
machine with ≈7 GiB more headroom than our VQ-2.4bpw build, no cluster, no patches,
stock mlx-lm.
Changelog
2026-10-01 — vq-skipzero: dead rows dropped
The vq-skipzero format drops the codes and scales of output rows whose teacher weights are about 1e-29; those rows output exact zeros, and in all 138 skip-zero modules (d8 K16384 and d4) the live rows are byte-identical to the previous revision (vqlab sz-check). Text weights go from 100.01 to 88.72 GiB (-11.29 GiB, 10.62%), and the full download shrinks by the same amount. This rung's expert codebooks are mostly d8, so it uses the d8 skip-zero kernels added to the bundled runtime today, which are byte-equal to the expanded form on synthetic tests at this rung's shapes, and the runtime also supports Knurlogic tensor-parallel splits of skip-zero modules. This revision is not KL-scored; generation was verified 2026-10-01 through Knurlogic split across two Macs (M3 Ultra 96 GB + M4 Max 128 GB) at 37.4 tok/s decode on a short prompt (not a paired benchmark).
2026-10-01 — runtime refresh; Knurlogic
The bundled model.py is updated to the current VQ runtime; weights are
unchanged. The run, multi-Mac, drafting and vision instructions now use
Knurlogic.
2026-09-20 — neighboring rungs corrected; this one is unchanged
The 2.4, 2.6 and 3.1 rungs were found to carry affine 3-bit on the expert modules of layers 57-59 while being described as uniform VQ, and have been rebuilt and re-measured; their rows in the KL table below now reflect the corrected weights. This rung is not affected — the v2 build served here covers all 60 layers — so its bytes and its numbers are unchanged.
2026-09-17 — the full Qwen3.5-397B-A17B VQ ladder on one KL instrument
KL to this family's own bf16 teacher (cached top-64), 12288 tokens, on all three house corpora — prose, public code and literary — every rung scored on the same cache and the same positions, through a streamed pass validated against a direct full-model forward to all printed decimals.
| build | GiB (text) | prose | code | literary | mean | top-1 |
|---|---|---|---|---|---|---|
| VQ-2.2bpw (this; v2, mixed) | 88.7 | 276.4 | 98.6 | 183.1 | 186.0 | 84.8% |
| VQ-2.4bpw | 95.7 | 230.4 | 87.7 | 134.1 | 150.7 | 86.2% |
| VQ-2.6bpw | 105.5 | 164.1 | 58.1 | 60.5 | 94.2 | 88.5% |
| spicyneuron 2.6bit (affine) | 120.6 | 332.1 | 98.4 | 241.0 | 223.8 | 83.5% |
| VQ-3.1bpw | 125.3 | 91.4 | 32.3 | 16.3 | 46.7 | 91.7% |
Millinats per token; lower is closer to bf16. Sizes are current; the VQ rows' KL and top-1 were scored on their full-row revisions, before the 2026-10-01 skip-zero rebuilds (which keep live rows byte-identical). These rows are VQ rungs only and are not comparable to the affine-comparator table further down, which was measured prose-only at 2048 tokens on an earlier instrument.
Rank by KL, not perplexity. On this family perplexity is not monotone in quality: rungs that are measurably closer to the teacher can read a higher perplexity, because a damaged model can score better than bf16 on a finite sample. KL on cached teacher logits does not have that failure mode.
2026-09-09 — runtime refresh
Runtime refresh. The bundled model.py is updated so downloaders run
exactly the code that was benchmarked; dense bundles now carry both runtimes.
- Faster prefill on affected geometries, measured per rung: a device-codebook kernel arm for large-codebook geometries (up to 1.46x on affected rungs), a ragged-subvector relaxation (up to 1.34x on affected rungs), a fused d8 arm, and a routing memo (≈2%). No blanket speedup is claimed across the lineup — gains apply only where the geometry engages the new paths.
- Output quality is unchanged: the kernel changes are bit-identical or 1-ULP-equivalent, and the routing memo is bit-identical (logits checksum verified).
- Speculative decoding (MTP): on repos that ship
mtp-head-q6.safetensors, the sidecar works with the exo fork branchmtp-stage1(github.com/noahzelezny/exo) — launch each node withexo --mtp(or setEXO_MTP=1). Note: with MTP enabled, exo serves requests sequentially (the batch engine has no MTP path), so leave it off for concurrent / multi-agent workloads.
2026-09 — bundle refresh
Bundle refresh.
model.py updated: refreshed VQ expert kernels
(verified equivalent; on the cluster the interconnect dominates, so
decode is unchanged — the refresh is for runtime consistency across the
lineup), a default buffer-cache ceiling so transient prefill
allocations return to the OS as they free, and dual-runtime loading.
Weights unchanged; prior revision pinnable.
Memory, measured externally (process RSS sampled at 5 Hz, M4 128 GB, weights ≈100 GiB on disk, full-row revision): 60.9 GiB peak for load + 2048-token prefill + 128-token decode — expert routing concentrates and untouched expert pages never load. Budget the weight size for worst-case varied traffic. Peak ≈ what routing touches, plus ≈3 GiB.
Measured results
See also the 2026-09-17 entry in the changelog above, which places this rung against the rest of its family on the three-corpus KL instrument. The table below is the earlier prose-only measurement and is not comparable to it row-for-row.
All numbers measured on this exact artifact, reproduced bit-identically ×2,
scored with an unmodified mlx-lm install. Read the whole row, not one cell:
| this model, v2 (88.7 GiB) | v1 (100.1 GiB) | spicyneuron 2.6bit (120.6 GiB) | VQ-2.4bpw (95.7 GiB) |
|
|---|---|---|---|---|
| wikitext perplexity (raw, prefix-8192) | 3.0591 | 3.1706 | 3.1843 | 2.7655 |
| code perplexity (mixed-language) | 2.6728 | 2.6988 | 2.6667 | 2.6383 |
Perplexities for this model and VQ-2.4bpw were scored on their full-row revisions, before the 2026-10-01 skip-zero rebuilds.
The honest trade: v2 improves on v1 by 3.5% on prose and 1.0% on code at
the same size and the same bits per weight. Against the closest community
quant it is better on prose (−3.9%) and better on code (+0.2% — within noise),
at 31.9 GiB smaller. Against VQ-2.4bpw it still gives up real quality on
both corpora.
What the size buys. A "128 GB" machine has 119.2 GiB of usable memory (vendors count in decimal GB; memory is allocated in binary GiB):
| build | resident at 8k context | free for KV cache, OS, everything else |
|---|---|---|
| this build | ≈90.5 GiB | ≈28.7 GiB |
VQ-2.4bpw |
≈95.7 GiB | ≈23.5 GiB |
That is the reason this build exists. VQ-2.4bpw is the better model and it
fits a 128 GB machine with very little room to spare. If your machine runs
VQ-2.4bpw comfortably at the context you need, use that one.
Speed. v2's 16,384-entry codebook no longer fits in Metal threadgroup memory, so it reads the codebook from device memory and runs roughly 20% slower than v1. We are not publishing throughput figures: repeat runs of the same artifact on the same machine varied more than the effect we would be reporting, and an unreliable number is worse than none. Measure on your own hardware.
Note the size does not buy speed in any case — this is an A17B MoE, so decode reads the same active experts per token as the larger builds. It buys residency, which is the table above.
Task benchmarks
All five models below — this repo's three VQ artifacts and the two community
comparators — were evaluated on the same harness, same settings, same
seeded items: lm-eval 0.4.12 driven by a layer-streaming loglikelihood
scorer (mlx-lm 0.31.3), 0-shot, first 1000 items per task, acc_norm
for HellaSwag/PIQA, acc for WinoGrande. Task numbers published elsewhere
come from a different pipeline and are not directly comparable, so the
comparator artifacts were re-evaluated here under identical conditions
rather than quoting their reported figures.
| model | size | HellaSwag | PIQA | WinoGrande |
|---|---|---|---|---|
| Qwen3.5-397B-A17B-VQ-2.2bpw (v1 weights) | 100.1 GiB | 0.861 | 0.841 | 0.787 |
| Qwen3.5-397B-A17B-VQ-2.4bpw | 108.0 GiB | 0.883 | 0.844 | 0.784 |
| spicyneuron 2.6bit | 120.6 GiB | 0.880 | 0.841 | 0.771 |
| Qwen3.5-397B-A17B-VQ-3.1bpw | 141.7 GiB | 0.903 | 0.840 | 0.780 |
| spicyneuron 3.5bit | 165.6 GiB | 0.904 | 0.846 | 0.767 |
Every model scored identical items, so differences are paired (McNemar exact test). HellaSwag reliably separates these quants and reproduces the perplexity ordering; PIQA and WinoGrande separate no pair at n=1000 and stand as integrity checks rather than rankings.
These rows were measured on v1's weights; v2 has not been re-evaluated on this harness, and the row above is labeled accordingly rather than reused for different weights. v2 improves on v1 on both perplexity corpora, but that is not a task-suite result and is not presented as one.
v1 was statistically indistinguishable from every larger model here on PIQA and WinoGrande; on HellaSwag it trails them by 2–4 points (paired p < 0.02) — the measured cost of the smallest size in this comparison. It is the accessibility artifact: the one that runs on a 128 GB Mac with real headroom.
These are 0-shot scores. Leaderboard conventions often use 10-shot HellaSwag / 5-shot WinoGrande, which run several points higher — compare against other 0-shot numbers only.
Speed
Measured 2026-10-01 on this revision.
Split across two Macs with Knurlogic (M3 Ultra 96 GB + M4 Max 128 GB, pipeline over Thunderbolt), against Qwen3.5-397B-A17B-MLX-2.6bit on the same two Macs: 3 pairs, a prompt of about 1,900 tokens, 128 generated tokens, no drafting.
| this model vs Qwen3.5-397B-A17B-MLX-2.6bit | |
|---|---|
| decode | 0.86× |
| prefill | 0.95× |
A two-Mac split's absolute speed depends on the link, so only the ratio is quoted. These ratios were measured on the full-row revision, before the 2026-10-01 skip-zero rebuild.
Run it
With Knurlogic, which works out the settings and checks the model fits before loading it:
pip install knurlogic
hf download TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.2bpw --local-dir ~/Knurlogic/Models/Qwen3.5-397B-A17B-VQ-2.2bpw
knurlogic serve ~/Knurlogic/Models/Qwen3.5-397B-A17B-VQ-2.2bpw
Chat at http://127.0.0.1:8080/, or point any OpenAI-, Anthropic- or
Ollama-compatible client at it (model local). knurlogic ui opens a page
with a Launch button for every model on disk instead.
At 102 GB on disk this needs a Mac with more memory than that, or two Macs:
install the same Knurlogic version and the model on each, run
knurlogic ui --host cluster on both, and launch it from the page with both
machines selected. Knurlogic splits the layers between them.
Images work through Knurlogic directly; mlx-vlm is not needed.
The VQ runtime ships inside this repo as model.py (declared by
model_file in config.json), and Knurlogic runs the file the model ships.
Without Knurlogic:
pip install 'mlx-lm>=0.31.3'
python -m mlx_lm generate \
--model TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.2bpw \
--prompt "Explain vector quantization briefly." \
--max-tokens 1000
Speculative decoding (MTP)
This repo includes mtp-head-q6.safetensors (5.4 GiB): the model's own
multi-token-prediction head, quantized. Stock loaders ignore it; it
costs nothing unless you opt in by name, and adds ≈5.6 GiB resident
when enabled.
Measured on this artifact: draft acceptance 0.72 (256 positions, control-ruled-out). Knurlogic drafts with it automatically (see below); VQLab's server also runs it single-box — this rung fits a 128 GB machine:
git clone https://github.com/noahzelezny/VQLab && cd VQLab
python3 -m venv .venv && source .venv/bin/activate
pip install .
python -m vqlab.cli serve \
--model TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.2bpw \
--sidecar mtp-head-q6.safetensors
Cluster speculative decoding is live: Knurlogic loads the shipped head and drafts automatically, on one Mac or split across several, with nothing to enable (outputs exactly the base model's via rejection sampling). Measured on this family's 2.6bpw rung on a 2-node exo pipeline: acceptance 0.85, throughput at parity with exo's stock decode — the head drafts well, but this family's stock pipeline decode does not degrade with generation length, so there is little for speculation to recover. Expect parity; per-rung numbers for this artifact are not yet measured.
Vision
The artifact includes the full 333-tensor vision tower at source precision
(0.85 GiB). mlx-lm is text-only for this architecture and ignores it;
Knurlogic loads it from this folder directly, without mlx-vlm.
The sizes quoted above are the download: they include this tower. Because
mlx-lm does not load it, resident memory runs ≈0.85 GiB below the disk
figure — the runtime tables report what was actually measured resident.
Tuning: prefill speed
mlx-lm prefills prompts in 512-token steps by default. This model is a
sparse MoE — a larger step puts more rows through each expert per call, which
this quantization format likes. The throughput figures below were measured
on v1; the knob and the memory costs apply unchanged to v2, the absolute
rates will be somewhat lower. Measured on an M4 Max 128 GB, 8k context:
--prefill-step-size |
prefill tok/s | peak memory |
|---|---|---|
| 512 (default) | ≈63–76 | 101.8 GiB |
| 1024 | ≈89 | 103.0 GiB |
| 2048 | ≈112–118 | 105.4 GiB |
| 4096 | ≈138–141 | 109.6 GiB |
Decode is unaffected (≈18–21 tok/s throughout) — it uses a different code path. The cost is memory: budget the peak above plus your KV cache. On a 128 GB machine this model has room for step 4096; leave headroom if you run long contexts. (Measured single-box; with more memory or a multi-Mac cluster the same knob applies with a higher ceiling.)
Note: perplexity is deterministic; wall-time figures are not, and will vary with whatever else your machine is doing.
Methodology
Mixed precision by layer sensitivity. Not all weights deserve the same bits. Attention, MoE routers, embeddings, and the output head stay at higher precision — they are a small fraction of the parameters but errors there propagate through every token. The MoE experts are ≈85% of the model and individually far more tolerant, so they absorb the aggressive quantization. A tail of later layers is also promoted above the expert baseline; measured layer-wise error showed depth matters, and the last layers repay the bits.
Vector quantization instead of scalar rounding — the part that is different. Scalar 2-bit gives each weight 4 rigid levels; over a group of 4 weights that is 256 fixed grid combinations. This build instead learns a codebook of joint 8-weight patterns and stores one index per group. Each 8-weight subvector stores one 14-bit index into a per-tensor 16,384-entry fp16 codebook. At the same bits, the codebook's entries sit where the weight distribution actually is, rather than on a uniform lattice — which is why this beats scalar quantization at matched size rather than merely matching it. Per-tensor codebooks, with an fp16 scale per (row, 64 weights), for 2.00 bits/weight stored in the expert region.
Codebooks are fit in pure weight space — k-means over the weight subvectors, no Hessian, no activation statistics, no calibration corpus. That is a deliberate choice: calibration-fitted methods we tested (GPTQ- and DWQ-style) reduced layer error while making end-to-end perplexity worse on this architecture, and they bias the result toward whatever text the calibration set contains. Weight-space fitting has no such domain preference.
Sub-byte bit-packing. Codes are packed into uint32 words (row-local, 32-code blocks) rather than padded to whole bytes, which is what makes the non-byte-aligned sizes possible at all. Packing is a pure representation change: the packed artifact's perplexities match its unpacked twin to four decimals of total negative log-likelihood on both corpora.
How it was evaluated. Perplexity on two corpora — raw wikitext
(prefix-8192) and a mixed-language code corpus — every number reproduced
bit-identically twice, scored with an unmodified mlx-lm. Two corpora
because this family shows real domain asymmetry: larger codebooks buy far
more on prose than on code, so a single-corpus number would misrepresent the
trade. Task-suite results are reported below, measured the same way.
Verification
Release gates passed on this artifact before upload: file, index and tokenizer checks, a verbatim match between the bundled runtime and its source, and a generation smoke through the shipping runtime on Apple Silicon. The upload path runs the gate itself and refuses to publish without it.
Limitations
- Code-heavy workloads measurably prefer the
VQ-2.4bpwbuild (+1.2% code perplexity here vs the community 2.6bit, +2.3% vs ourVQ-2.4bpwbuild). - This is a thinking model (Qwen3.5 family): it spends tokens reasoning
before answering. Budget
max_tokensaccordingly. - To split this model across two or more Macs, use Knurlogic (see Run it). Single-box users are unaffected.
Acknowledgment
spicyneuron's 397B quants are what made this model runnable on my hardware in the first place — they were the artifacts that fit when nothing else did, and they were the reference this work was measured against throughout. This release is offered in that same spirit: the full method, the experiments that failed as well as the ones that worked, and comparator numbers re-measured on one harness so the claims can be checked rather than taken on trust.
Paper
The method, the full model ladder, the negative results, and the measurement rules behind every number here: Data-Free Vector Quantization Beats Affine Quantization at Matched Bytes Below 6 Bits (CC BY 4.0) · code: VQLab · web version: Space
Support this work
VQLab and these artifacts are built and released independently — the fits, the measurement harness and the published rungs are one person's compute and time. If they are useful to you:
- Sponsor: github.com/sponsors/noahzelezny
- Contract work: available for quantization and on-device inference work on Apple silicon — custom rungs, per-layer allocation for your model, or getting a checkpoint to run well on a Mac. Contact: [email protected]
Provenance
Base model: Qwen/Qwen3.5-397B-A17B — Apache-2.0. This is a quantized derivative and inherits that license; using it means accepting the base model's terms. Quantization: TheDrainFlorist, 2026.
Built with MLX and VQLab.