125.3 GiB text weights — the quality build. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 131.6 GiB. A vector-quantized build of Qwen3.5-397B-A17B for machines with memory to spend: the strongest quantization we know how to make of this model at this size, on stock mlx-lm, no patches. 12.43% of the 64-element weight groups in the Qwen3.5-397B teacher sit in output rows whose weights are about 1e-29. The vq-skipzero format drops those rows' codes and scales on disk and in memory; they output exact zeros, and the live rows are byte-identical to the previous revision (vqlab sz-check, every module). Text weights go from 141.71 to 125.31…
Open weights
apache-2.0
64.5B parameters
262,144 tokens
mlx
88.7 GiB text weights — the accessibility build, the roomiest fit on a 128 GB Mac. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 95.0 GiB. (v2, mixed geometry.) v2 — updated 2026-08-22. This repository now serves a rebuilt artifact at the same size and the same bits per weight, with a different codebook geometry that measures better on both perplexity corpora. v1's numbers are kept below rather than quietly overwritten, and v1's bytes remain downloadable by pinning the previous revision: A vector-quantized build of Qwen3.5-397B-A17B running? 88.7 GiB text weights — it runs on a single 128 GB Apple Silicon machine with ≈7 GiB more…
Open weights
apache-2.0
62.6B parameters
262,144 tokens
mlx
95.7 GiB text weights — the daily driver, runs on a single 128 GB Mac. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 101.9 GiB. A vector-quantized build of Qwen3.5-397B-A17B that fits and generates on one 128 GB Apple Silicon machine — no cluster, no patches, stock mlx-lm. 12.43% of the 64-element weight groups in the Qwen3.5-397B teacher sit in output rows whose weights are about 1e-29. The vq-skipzero format drops those rows' codes and scales on disk and in memory; they output exact zeros, and the live rows are byte-identical to the previous revision (vqlab sz-check, every module). Text weights go from 107.96 to 95.66 GiB (-12.30…
Open weights
apache-2.0
120.2B parameters
262,144 tokens
mlx
105.5 GiB text weights — the balanced build. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 111.8 GiB. It lands in the same size class as the strongest community quant at this rate; the table below is the comparison. A vector-quantized build of Qwen3.5-397B-A17B for Apple Silicon. Stock mlx-lm, no patches — the VQ runtime ships inside the checkpoint as model.py. It sits between the VQ-2.4bpw daily driver and the VQ-3bpw quality build, and needs the same hardware class as the latter. 12.43% of the 64-element weight groups in the Qwen3.5-397B teacher sit in output rows whose weights are about 1e-29. The vq-skipzero format drops those…
Open weights
apache-2.0
59.2B parameters
262,144 tokens
mlx