SAVRN
Search Contact SAVRN

Qwen3.5-397B-A17B-VQ-2.6bpw · Model Card

Qwen3.5-397B-A17B-VQ-2.6bpw: Model Card

Written by Noah Zelezny, published under apache-2.0, revision 5ff492af70af, read 2026-10-02. Shown as written; SAVRN's own facts about this model are on its page.

105.5 GiB text weights — the balanced build. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 111.8 GiB.

It lands in the same size class as the strongest community quant at this rate; the table below is the comparison. A vector-quantized build of Qwen3.5-397B-A17B for Apple Silicon. Stock mlx-lm, no patches — the VQ runtime ships inside the checkpoint as model.py. It sits between the VQ-2.4bpw daily driver and the VQ-3bpw quality build, and needs the same hardware class as the latter.

Changelog

2026-10-01 — vq-skipzero: dead rows dropped

12.43% of the 64-element weight groups in the Qwen3.5-397B teacher sit in output rows whose weights are about 1e-29. The vq-skipzero format drops those rows' codes and scales on disk and in memory; they output exact zeros, and the live rows are byte-identical to the previous revision (vqlab sz-check, every module). Text weights go from 119.21 to 105.54 GiB (-13.67 GiB), and resident memory drops by about the same amount. Paired 3-corpus KL against the previous revision (12288 tokens each, 2026-09-29): no corpus differs at |t|>2.

2026-10-01 — runtime refresh; Knurlogic

The bundled model.py is updated to the current VQ runtime; weights are unchanged. The run, multi-Mac, drafting and vision instructions now use Knurlogic.

2026-09-20 — correction: layers 57-59 were affine, and now are not

What was wrong. This rung was described as a uniform VQ build, but the expert modules of layers 57, 58 and 59 — 9 of 180 — were never fitted and fell through to the base affine 3-bit, group 64. The fit had been invoked over layers 0-56 of a 60-layer model, and the off-by-three then traveled as a constant into every downstream tool. Those three layers therefore carried more bits than the rung's nominal rate, not fewer, so the published quality numbers were mildly optimistic.

What changed. The 9 modules are refitted at this rung's own geometry from the bf16 teacher, so the build is now genuinely uniform across all 60 layers. It is slightly smaller and slightly worse, and both are expected: the correction removes bits from late layers, where they do the most good.

The KL table below has been re-measured on the corrected weights — same instrument, same teacher cache, same positions as before. Sizes on this card are now stated consistently as text weights (the number that governs whether it runs); the vision tower (+0.85 GiB) and the optional MTP draft head (+5.4 GiB) are real extra bytes and are called out wherever a size appears. Earlier revisions of this card mixed those three totals in one column.

This is a correction to the numbers, not to how the model is used: the runtime, the config surface, the vision tower and the MTP head are all unchanged, and text and image generation were both re-verified on a 2-node exo pipeline after the rebuild.

2026-09-17 — the full Qwen3.5-397B-A17B VQ ladder on one KL instrument

KL to this family's own bf16 teacher (cached top-64), 12288 tokens, on all three house corpora — prose, public code and literary — every rung scored on the same cache and the same positions, through a streamed pass validated against a direct full-model forward to all printed decimals.

build GiB (text) prose code literary mean top-1
VQ-2.2bpw (v2, mixed) 100.0 276.4 98.6 183.1 186.0 84.8%
VQ-2.4bpw 95.7 230.4 87.7 134.1 150.7 86.2%
VQ-2.6bpw (this) 105.5 164.1 58.1 60.5 94.2 88.5%
spicyneuron 2.6bit (affine) 120.6 332.1 98.4 241.0 223.8 83.5%
VQ-3.1bpw 125.3 91.4 32.3 16.3 46.7 91.7%

Millinats per token; lower is closer to bf16. The KL and top-1 rows were scored on the full-row revisions; sizes are the current vq-skipzero ones. These rows are VQ rungs only and are not comparable to the affine-comparator table further down, which was measured prose-only at 2048 tokens on an earlier instrument.

Rank by KL, not perplexity. On this family perplexity is not monotone in quality: rungs that are measurably closer to the teacher can read a higher perplexity, because a damaged model can score better than bf16 on a finite sample. KL on cached teacher logits does not have that failure mode.

2026-09-09 — runtime refresh

Runtime refresh. The bundled model.py is updated so downloaders run exactly the code that was benchmarked; dense bundles now carry both runtimes.

  • Faster prefill on affected geometries, measured per rung: a device-codebook kernel arm for large-codebook geometries (up to 1.46x on affected rungs), a ragged-subvector relaxation (up to 1.34x on affected rungs), a fused d8 arm, and a routing memo (≈2%). No blanket speedup is claimed across the lineup — gains apply only where the geometry engages the new paths.
  • Output quality is unchanged: the kernel changes are bit-identical or 1-ULP-equivalent, and the routing memo is bit-identical (logits checksum verified).
  • Speculative decoding (MTP): on repos that ship mtp-head-q6.safetensors, the sidecar works with the exo fork branch mtp-stage1 (github.com/noahzelezny/exo) — launch each node with exo --mtp (or set EXO_MTP=1). Note: with MTP enabled, exo serves requests sequentially (the batch engine has no MTP path), so leave it off for concurrent / multi-agent workloads.

2026-09 — bundle refresh (runtime update landed)

Bundle refresh (runtime update landed).

model.py updated: refreshed VQ expert kernels (verified equivalent), a default buffer-cache ceiling (VQLAB_CACHE_LIMIT_GB overrides, 0 disables) so transient prefill allocations return to the OS as they free — long-prompt peak stays near resident instead of a multiple of it — and dual-runtime loading. Weights unchanged; prior revision pinnable. This rung exceeds the single-box gate bar, so the update shipped through the release gate's cluster smoke: a real generation on a 2-node exo pipeline, with the peer rank's copy identity-checked before the generation counted.

Measured results

See also the 2026-09-17 entry in the changelog above, which places this rung against the rest of its family on the three-corpus KL instrument. The table below is the earlier prose-only measurement and is not comparable to it row-for-row.

Scored on this exact artifact with an unmodified mlx-lm install, on the same two corpora and the same harness as every comparator below:

this model (105.5 GiB text) spicyneuron 2.6bit (120.6 GiB text)
wikitext perplexity (raw, prefix-8192)* 2.5634 3.1843
code perplexity (mixed-language)* 2.6123 2.6667

19.5% better on prose and 2.0% better on code, at 15.1 GiB less on disk.

*wikitext/code perplexity here is the older external-corpus instrument, measured pre-correction; the KL table above is the current ranking, re-measured on the corrected weights.

Both margins are large relative to the fit-to-fit noise we can measure: 24x and 3.1x respectively. One caveat on those multiples, because it is the kind of thing that is easy to leave out — the noise floor they are quoted against was measured at a different codebook size (256, not 512). No floor has been measured at this geometry. Floors in this project have widened every time they were measured more carefully, so treat 24x as "comfortably real" rather than as a precise figure. The prose gap is not in doubt at any plausible floor; the code gap is the smaller of the two and would be the first to become uninteresting if this geometry's floor turned out to be wide.

The size comparison, measured. This artifact carries the full vision tower at bf16 (333 tensors, 0.849 GiB); the spicyneuron builds are text-only and carry none. Comparing text weights to text weights: this build is 105.5 GiB against their 120.6 — 15.1 GiB smaller, and ahead on all three corpora of the KL table above. Before the 2026-09-20 geometry correction it was 121.4 against their 120.6 (+0.8 GiB), so the correction moved this rung from slightly larger to slightly smaller than the comparator; the 2026-10-01 vq-skipzero revision removed another 13.7 GiB.

Sizes on this card are model weights, including the vision tower. The repo also ships an optional mtp-head-q6.safetensors drafting sidecar (5.41 GiB) which mlx-lm does not load; a full snapshot_download fetches it, so budget for it separately.

Task benchmarks

Not yet measured on this artifact. Two of the siblings above carry HellaSwag/PIQA/WinoGrande numbers; this build has not been run through that harness, and reporting a sibling's task scores here would be exactly the substitution this project refuses to make. They will be added once the suite has been run under the same harness (lm-eval 0.4.12, layer-streaming loglikelihood scorer, 0-shot, first 1000 items per task).

Hardware

This build has not been verified on a single 128 GB machine. Use either:

  • a single Apple Silicon machine with ≥ 192 GB unified memory, or
  • two or more Macs with Knurlogic (see Run it). This rung was verified on a 96 GB + 128 GB exo cluster over Thunderbolt.

Verified serving on a 2-node ring: placed and serving in 99 seconds, 800-token coherent generation, and three graded known-answer probes returned correct. No throughput figure is published here — we measured placement and correctness, not tokens per second, and an unmeasured number is worse than none.

This artifact has not been verified single-node. Any re-verification has to be a 2-node ring.

Speed

Measured 2026-10-01 on this revision.

Split across two Macs with Knurlogic (M3 Ultra 96 GB + M4 Max 128 GB, pipeline over Thunderbolt), against Qwen3.5-397B-A17B-MLX-2.6bit on the same two Macs: 3 pairs, a prompt of about 1,900 tokens, 128 generated tokens, no drafting.

this model vs Qwen3.5-397B-A17B-MLX-2.6bit
decode 0.83×
prefill 0.88×

A two-Mac split's absolute speed depends on the link, so only the ratio is quoted.

The ratios above were measured on the previous full-row revision. Generation on this revision was verified 2026-10-01 on the same two Macs at 38.2 tok/s decode (short prompt, 8-40 tokens; not a paired benchmark).

Run it

With Knurlogic, which works out the settings and checks the model fits before loading it:

pip install knurlogic
hf download TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.6bpw --local-dir ~/Knurlogic/Models/Qwen3.5-397B-A17B-VQ-2.6bpw
knurlogic serve ~/Knurlogic/Models/Qwen3.5-397B-A17B-VQ-2.6bpw

Chat at http://127.0.0.1:8080/, or point any OpenAI-, Anthropic- or Ollama-compatible client at it (model local). knurlogic ui opens a page with a Launch button for every model on disk instead.

At 120 GB on disk this needs a Mac with more memory than that, or two Macs: install the same Knurlogic version and the model on each, run knurlogic ui --host cluster on both, and launch it from the page with both machines selected. Knurlogic splits the layers between them.

Images work through Knurlogic directly; mlx-vlm is not needed.

The VQ runtime ships inside this repo as model.py (declared by model_file in config.json), and Knurlogic runs the file the model ships.

Without Knurlogic:

pip install 'mlx-lm>=0.31.3'
python -m mlx_lm generate \
  --model TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.6bpw \
  --prompt "Explain vector quantization briefly." \
  --max-tokens 1000

Speculative decoding (MTP)

mtp-head-q6.safetensors (5.4 GiB) is the model's own MTP draft head — the same file validated on the 2.2bpw rung (acceptance 0.72, single-box via vqlab serve --sidecar). Cluster speculative decoding is live: Knurlogic loads the shipped head and drafts automatically, on one Mac or split across several, with nothing to enable (outputs exactly the base model's via rejection sampling). Measured on this family's 2.6bpw rung on a 2-node exo pipeline: acceptance 0.85, throughput at parity with exo's stock decode — the head drafts well, but this family's stock pipeline decode does not degrade with generation length, so there is little for speculation to recover. Expect parity.

Vision

The artifact includes the full 333-tensor vision tower at source precision (0.85 GiB). mlx-lm is text-only for this architecture and ignores it; Knurlogic loads it from this folder directly, without mlx-vlm.

The sizes quoted above are the download: they include this tower. Because mlx-lm does not load it, resident memory runs ≈0.85 GiB below the disk figure.

Siblings

All from the same skeleton and recipe, all scored the same way:

size wikitext code needs
VQ-2.2bpw 100.0 GiB 3.0591 2.6728 128 GB Mac, roomy
VQ-2.4bpw 95.7 GiB 2.7655 2.6383 128 GB Mac, tight
VQ-2.6bpw (this build) 105.5 GiB 2.5634 2.6123 ≥192 GB or cluster
VQ-3.1bpw 125.3 GiB 2.3410 2.5963 ≥192 GB or cluster

If you can run this build you can run VQ-3bpw, which is better on both corpora for 19.8 GiB more. Take this one if those gigabytes are worth more to you than the quality difference — on a 192 GB machine it leaves roughly 84 GiB free against the 3bpw build's 64.

Methodology

Mixed precision by layer sensitivity. Attention, MoE routers, embeddings and the output head stay at higher precision — a small fraction of the parameters, but errors there propagate through every token. The MoE experts are ≈85% of the model and individually far more tolerant, so they absorb the aggressive quantization.

Vector quantization instead of scalar rounding — the part that is different. Scalar 2-bit gives each weight 4 rigid levels; over a group of 4 weights that is 256 fixed grid combinations. This build learns a codebook of joint 4-weight patterns and stores one index per group. Each 4-weight subvector stores one 9-bit index into a per-tensor 512-entry fp16 codebook, with an fp16 scale per (row, 64 weights) — 2.5 bits/weight stored in the expert region. At the same bits the codebook's entries sit where the weight distribution actually is, rather than on a uniform lattice.

Every expert tensor uses this one geometry; there is no mixed allocation and no per-layer schedule. Flat rungs are the reference points in this lineup because no mixed-allocation build we measured beat the flat rung at its own size.

Codebooks are fit in pure weight space — k-means over the weight subvectors, no Hessian, no activation statistics, no calibration corpus. Calibration-fitted methods we tested (GPTQ- and DWQ-style) reduced layer error while making end-to-end perplexity worse on this architecture, and they bias the result toward whatever text the calibration set contains.

The fit is not seeded. k-means draws an unseeded subsample, so this artifact is reproducible in recipe and geometry but not bit-for-bit. That is why margins here are quoted against a measured fit-to-fit floor rather than against a repeated build.

Sub-byte bit-packing. Codes are packed into uint32 words (row-local, 32-code blocks) rather than padded to whole bytes, which is what makes the non-byte-aligned size possible. Packing is a pure representation change.

How it was evaluated. Perplexity on two corpora — raw wikitext (prefix-8192) and a mixed-language code corpus — scored with an unmodified mlx-lm on the same harness used for every comparator here. Two corpora because this family shows real domain asymmetry: larger codebooks buy far more on prose than on code, so a single-corpus number would misrepresent the trade.

Verification

Release gates passed on this artifact before upload: file, index and tokenizer checks, a verbatim match between the bundled runtime and its source, and a generation smoke through the shipping runtime on Apple Silicon. The upload path runs the gate itself and refuses to publish without it.

Limitations

  • No throughput measurement — see Hardware.
  • Perplexities are corpus-specific. Compare only against models scored on the same files, never across harnesses.
  • This is a thinking model: it spends tokens reasoning before answering. Budget max_tokens accordingly.

Acknowledgment

spicyneuron's 397B quants are what made this model runnable on my hardware in the first place, and they were the reference this work was measured against throughout. This release is offered in that spirit: the method, the failures as well as the wins, and comparator numbers re-measured on one harness so the claims can be checked rather than taken on trust.

Paper

The method, the full model ladder, the negative results, and the measurement rules behind every number here: Data-Free Vector Quantization Beats Affine Quantization at Matched Bytes Below 6 Bits (CC BY 4.0) · code: VQLab · web version: Space

Support this work

VQLab and these artifacts are built and released independently — the fits, the measurement harness and the published rungs are one person's compute and time. If they are useful to you:

  • Sponsor: github.com/sponsors/noahzelezny
  • Contract work: available for quantization and on-device inference work on Apple silicon — custom rungs, per-layer allocation for your model, or getting a checkpoint to run well on a Mac. Contact: [email protected]

Provenance

Base model: Qwen/Qwen3.5-397B-A17B — Apache-2.0. This is a quantized derivative and inherits that license; using it means accepting the base model's terms. Quantization: TheDrainFlorist, 2026.

Built with MLX and VQLab.