# Qwen3.5-397B-A17B-VQ-2.2bpw by Noah Zelezny: Open Model
Source: https://savrn.com/models/qwen3-5-397b-a17b-vq-2-2bpw
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve Qwen3.5-397B-A17B-VQ-2.2bpw (62.6B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 125.1 GB | 150.2 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 · [1x MI355X](https://savrn.com/ai-index/pricing/gpus/mi355x) $2.59 |
| 8-bit | 62.6 GB | 75.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 4-bit | 31.3 GB | 37.5 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 7, 2026.

[Qwen3.5-397B-A17B-VQ-2.2bpw on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/qwen3-5-397b-a17b-vq-2-2bpw/gpus)

## Model Card

By Noah Zelezny, published under apache-2.0, revision f10167305589.

88.7 GiB text weights — the accessibility build, the roomiest fit on a 128 GB Mac. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 95.0 GiB. (v2, mixed geometry.)

v2 — updated 2026-08-22. This repository now serves a rebuilt artifact at the same size and the same bits per weight, with a different codebook geometry that measures better on both perplexity corpora. v1's numbers are kept below rather than quietly overwritten, and v1's bytes remain downloadable by pinning the previous revision:

```
snapshot_download("TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.2bpw",
                  revision="4554635165011f67e8166fd94d4bcc8cbf91401c")  # v1
```

A vector-quantized build of [Qwen3.5-397B-A17B](https://huggingface.co/Qwen/Qwen3.5-397B-A17B) built to answer one question: how small can a 397B get and still be worth running? 88.7 GiB text weights — it runs on a single 128 GB Apple Silicon machine with ≈7 GiB more headroom than our VQ-2.4bpw build, no cluster, no patches, stock mlx-lm.

### Changelog

#### 2026-10-01 — vq-skipzero: dead rows dropped

[Read the full model card (2,914 words)](https://savrn.com/models/qwen3-5-397b-a17b-vq-2-2bpw/card)

## Configuration

Architecture

Qwen3_5MoeForConditionalGeneration

Context length (tokens)

262,144

Layers

60

Hidden size

4,096

Attention heads

32

Key/value heads

2

Head dimension

256

Vocabulary size

248,320

Experts

512

Experts active per token

10

Model type

qwen3_5_moe

## Identity and Version

Repository

TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.2bpw

Publisher

Noah Zelezny

Task

Text generation

Modality

Text

Library

mlx

Parameters

62.6B parameters

Languages

en

Revision

f101673055892f18fd0645c3f9af8e1c5f7093a2

First published

2026-08-16

Last updated

2026-10-02

## Files and Weights

44 files, 107.8 GB in total. The weights are 29 files totalling 102.0 GB in safetensors.

Weights29 files · 102.0 GB

Configuration8 files · 714.9 KB

Tokenizer2 files · 20.0 MB

Documentation1 file · 20.4 KB

Other3 files · 5.8 GB

Repository1 file · 1.6 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| model-00001-of-00027.safetensors | Weights | 2.2 GB | 4c9d468f14fd |
| model-00002-of-00027.safetensors | Weights | 2.3 GB | 53a0a50986d8 |
| model-00003-of-00027.safetensors | Weights | 2.9 GB | 7931585f0181 |
| model-00004-of-00027.safetensors | Weights | 3.2 GB | 830f8012d9a8 |
| model-00005-of-00027.safetensors | Weights | 3.0 GB | f41620a614e8 |
| model-00006-of-00027.safetensors | Weights | 3.1 GB | aef6a858297c |
| model-00007-of-00027.safetensors | Weights | 3.5 GB | d22ab12bda77 |
| model-00008-of-00027.safetensors | Weights | 3.4 GB | 4dbb6ccfcd44 |
| model-00009-of-00027.safetensors | Weights | 3.5 GB | e9d20b1fd94c |
| model-00010-of-00027.safetensors | Weights | 3.7 GB | 90445e50e95c |
| model-00011-of-00027.safetensors | Weights | 3.7 GB | d5a030ac02e8 |
| model-00012-of-00027.safetensors | Weights | 3.7 GB | 003895d9f2a3 |
| model-00013-of-00027.safetensors | Weights | 4.0 GB | 24671c33a20d |
| model-00014-of-00027.safetensors | Weights | 3.9 GB | 9871794b6b74 |
| model-00015-of-00027.safetensors | Weights | 3.9 GB | aec2476a3a9b |
| model-00016-of-00027.safetensors | Weights | 4.0 GB | d2b1f4f57771 |
| model-00017-of-00027.safetensors | Weights | 4.0 GB | 473f1c4c4b77 |
| model-00018-of-00027.safetensors | Weights | 4.0 GB | de0b845182b3 |
| model-00019-of-00027.safetensors | Weights | 4.2 GB | cf5298e7cfb2 |
| model-00020-of-00027.safetensors | Weights | 4.0 GB | f7653046f408 |
| model-00021-of-00027.safetensors | Weights | 4.1 GB | cc39e393a478 |
| model-00022-of-00027.safetensors | Weights | 4.1 GB | 4b3ec939a385 |
| model-00023-of-00027.safetensors | Weights | 3.7 GB | d1aa71049342 |
| model-00024-of-00027.safetensors | Weights | 3.7 GB | 344a4a809cf4 |
| model-00025-of-00027.safetensors | Weights | 3.5 GB | 3cca0405ecc8 |
| model-00026-of-00027.safetensors | Weights | 3.7 GB | 4b92275e0669 |
| model-00027-of-00027.safetensors | Weights | 2.2 GB | ae7440e308ff |
| model-vision-graft.safetensors | Weights | 912.1 MB | b47e83150554 |
| mtp-head-q6.safetensors | Weights | 5.8 GB | 1563d297c7bc |
| config.json | Configuration | 148.7 KB | — |
| generation_config.json | Configuration | 244 B | — |
| model.py | Configuration | 262.3 KB | — |
| model.safetensors.index.json | Configuration | 288.6 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| skipzero_load.py | Configuration | 6.3 KB | — |
| video_preprocessor_config.json | Configuration | 385 B | — |
| vqlab_provenance.json | Configuration | 8.1 KB | — |
| README.md | Documentation | 20.4 KB | — |
| chat_template.jinja | Other | 7.8 KB | — |
| mtp-head-q6.safetensors.fp32norms-bak | Other | 5.8 GB | a5cad76f094e |
| vqlab_provenance.history.jsonl | Other | 52.8 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 20.0 MB | 06b9509352d2 |
| tokenizer_config.json | Tokenizer | 1.2 KB | — |

## License and Download

License

apache-2.0

Access

Open weights, no gate

Download size

102.0 GB

[Download from Noah Zelezny](https://huggingface.co/TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.2bpw)

Released by Noah Zelezny through its official repository on Hugging Face. [Read the license](https://www.apache.org/licenses/LICENSE-2.0).

## Built From

- Derived from Qwen/Qwen3.5-397B-A17B
- Quantized from Qwen/Qwen3.5-397B-A17B

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 102.0 GB |
| 16-bit | 125.1 GB |
| 8-bit | 62.6 GB |
| 4-bit | 31.3 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About Qwen3.5-397B-A17B-VQ-2.2bpw

### How much GPU memory does Qwen3.5-397B-A17B-VQ-2.2bpw need?

About 150.2 GB at 16-bit and 37.5 GB at 4-bit: the weights (62.6B parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run Qwen3.5-397B-A17B-VQ-2.2bpw on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### Can I use Qwen3.5-397B-A17B-VQ-2.2bpw commercially?

Yes. Qwen3.5-397B-A17B-VQ-2.2bpw is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

### What is Qwen3.5-397B-A17B-VQ-2.2bpw's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

## Similar Models

Model · Text generation

### [less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-g-50](https://savrn.com/models/less-is-moe-qwen3-5-122b-a10b-gpqa-main-64-intdim-g-50)

[XINKAI ZOU](https://savrn.com/model-publishers/jayzou3773)

This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,632 tokens. The source checkpoint was loaded and pruned in BF16. The source-row selection hash is…

Open weights apache-2.0 64.1B parameters 262,144 tokens

[View model](https://savrn.com/models/less-is-moe-qwen3-5-122b-a10b-gpqa-main-64-intdim-g-50)

Model · Text generation

### [less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-l-50](https://savrn.com/models/less-is-moe-qwen3-5-122b-a10b-gpqa-main-64-intdim-l-50)

[XINKAI ZOU](https://savrn.com/model-publishers/jayzou3773)

This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,632 tokens. The source checkpoint was loaded and pruned in BF16. The source-row selection hash is…

Open weights apache-2.0 64.1B parameters 262,144 tokens

[View model](https://savrn.com/models/less-is-moe-qwen3-5-122b-a10b-gpqa-main-64-intdim-l-50)

Model · Text generation

### [less-is-moe-qwen3.5-122b-a10b-gpqa-main-64-intdim-e-50](https://savrn.com/models/less-is-moe-qwen3-5-122b-a10b-gpqa-main-64-intdim-e-50)

[XINKAI ZOU](https://savrn.com/model-publishers/jayzou3773)

This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,632 tokens. The source checkpoint was loaded and pruned in BF16. The source-row selection hash is…

Open weights apache-2.0 64.1B parameters 262,144 tokens

[View model](https://savrn.com/models/less-is-moe-qwen3-5-122b-a10b-gpqa-main-64-intdim-e-50)

Model · Text generation

### [Qwen3.5-397B-A17B-VQ-3.1bpw](https://savrn.com/models/qwen3-5-397b-a17b-vq-3-1bpw)

[Noah Zelezny](https://savrn.com/model-publishers/thedrainflorist)

125.3 GiB text weights — the quality build. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 131.6 GiB. A vector-quantized build of Qwen3.5-397B-A17B for machines with memory to spend: the strongest quantization we know how to make of this model at this size, on stock mlx-lm, no patches. 12.43% of the 64-element weight groups in the Qwen3.5-397B teacher sit in output rows whose weights are about 1e-29. The vq-skipzero format drops those rows' codes and scales on disk and in memory; they output exact zeros, and the live rows are byte-identical to the previous revision (vqlab sz-check, every module). Text weights go from 141.71 to 125.31…

Open weights apache-2.0 64.5B parameters 262,144 tokens mlx

[View model](https://savrn.com/models/qwen3-5-397b-a17b-vq-3-1bpw)

Model · Text generation

### [Qwen3.5-122B-A10B-NVFP4](https://savrn.com/models/qwen3-5-122b-a10b-nvfp4)

[NVIDIA](https://savrn.com/model-publishers/nvidia)

The NVIDIA Qwen3.5-122B-A10B-NVFP4 model is the quantized version of Alibaba's Qwen3.5-122B-A10B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.5-122B-A10B NVFP4 model is quantized with Model Optimizer. This model is ready for commercial/non-commercial use. This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case; see link to Non-NVIDIA (Qwen3.5-122B-A10B) Model Card from Alibaba. Global Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems…

Open weights apache-2.0 64.6B parameters 262,144 tokens Model Optimizer

[View model](https://savrn.com/models/qwen3-5-122b-a10b-nvfp4)

Model · Text generation

### [less-is-moe-gpt-oss-120b-gpqa-main-64-intdim-e-50](https://savrn.com/models/less-is-moe-gpt-oss-120b-gpqa-main-64-intdim-e-50)

[XINKAI ZOU](https://savrn.com/model-publishers/jayzou3773)

This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,511 tokens. The source MXFP4 checkpoint was explicitly dequantized to BF16 before scoring and pruning. The source-row selection hash is…

Open weights apache-2.0 59.5B parameters 131,072 tokens

[View model](https://savrn.com/models/less-is-moe-gpt-oss-120b-gpqa-main-64-intdim-e-50)

## Noah Zelezny

[All models and datasets](https://savrn.com/model-publishers/thedrainflorist)

## Versions

- [f10167305589](https://savrn.com/models/qwen3-5-397b-a17b-vq-2-2bpw/versions/f10167305589) · current 2026-10-02

## Explore More

- [All text generation models](https://savrn.com/models/tasks/text-generation)
- [All models under apache-2.0](https://savrn.com/models/licenses/apache-2-0)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-10-02.
- [Hugging Face record](https://huggingface.co/TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.2bpw)
- [How the hub is built](https://savrn.com/model-hub/methodology)
