# Qwen3.5-397B-A17B-VQ-2.4bpw by Noah Zelezny: Open Model
Source: https://savrn.com/models/qwen3-5-397b-a17b-vq-2-4bpw
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve Qwen3.5-397B-A17B-VQ-2.4bpw (120.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 240.4 GB | 288.4 GB | 2x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $3.70 | [2x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $4.00 · [2x MI355X](https://savrn.com/ai-index/pricing/gpus/mi355x) $5.18 |
| 8-bit | 120.2 GB | 144.2 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 · [1x MI355X](https://savrn.com/ai-index/pricing/gpus/mi355x) $2.59 |
| 4-bit | 60.1 GB | 72.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 7, 2026.

[Qwen3.5-397B-A17B-VQ-2.4bpw on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/qwen3-5-397b-a17b-vq-2-4bpw/gpus)

## Model Card

By Noah Zelezny, published under apache-2.0, revision 34136c9eaf83.

95.7 GiB text weights — the daily driver, runs on a single 128 GB Mac. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 101.9 GiB.

A vector-quantized build of [Qwen3.5-397B-A17B](https://huggingface.co/Qwen/Qwen3.5-397B-A17B) that fits and generates on one 128 GB Apple Silicon machine — no cluster, no patches, stock mlx-lm.

### Changelog

#### 2026-10-01 — vq-skipzero: dead rows dropped

[Read the full model card (2,270 words)](https://savrn.com/models/qwen3-5-397b-a17b-vq-2-4bpw/card)

## Configuration

Architecture

Qwen3_5MoeForConditionalGeneration

Context length (tokens)

262,144

Layers

60

Hidden size

4,096

Attention heads

32

Key/value heads

2

Head dimension

256

Vocabulary size

248,320

Experts

512

Experts active per token

10

Model type

qwen3_5_moe

## Identity and Version

Repository

TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.4bpw

Publisher

Noah Zelezny

Task

Text generation

Modality

Text

Library

mlx

Parameters

120.2B parameters

Languages

en

Revision

34136c9eaf836441dde5595ba5703d26d76a17aa

First published

2026-08-16

Last updated

2026-10-02

## Files and Weights

44 files, 115.3 GB in total. The weights are 29 files totalling 109.4 GB in safetensors.

Weights29 files · 109.4 GB

Configuration8 files · 711.9 KB

Tokenizer2 files · 20.0 MB

Documentation1 file · 15.9 KB

Other3 files · 5.8 GB

Repository1 file · 1.6 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| model-00001-of-00027.safetensors | Weights | 2.7 GB | f2126104e5c4 |
| model-00002-of-00027.safetensors | Weights | 2.6 GB | 3bc4dbfaa827 |
| model-00003-of-00027.safetensors | Weights | 3.2 GB | c54d3ece77cb |
| model-00004-of-00027.safetensors | Weights | 3.6 GB | 792a3a05d1fa |
| model-00005-of-00027.safetensors | Weights | 3.4 GB | b8410f8d61ba |
| model-00006-of-00027.safetensors | Weights | 3.5 GB | 7898332d6c83 |
| model-00007-of-00027.safetensors | Weights | 3.9 GB | 5f28af41e4c7 |
| model-00008-of-00027.safetensors | Weights | 3.8 GB | 8eece7466147 |
| model-00009-of-00027.safetensors | Weights | 3.9 GB | 1709bb80eb39 |
| model-00010-of-00027.safetensors | Weights | 4.2 GB | 4ca01a8105d3 |
| model-00011-of-00027.safetensors | Weights | 4.1 GB | 970b2b2f8cab |
| model-00012-of-00027.safetensors | Weights | 4.2 GB | c8ccda9bf991 |
| model-00013-of-00027.safetensors | Weights | 4.3 GB | ccbfda45cff0 |
| model-00014-of-00027.safetensors | Weights | 4.3 GB | e74624b50012 |
| model-00015-of-00027.safetensors | Weights | 4.3 GB | baa8cd774199 |
| model-00016-of-00027.safetensors | Weights | 4.4 GB | 14cbcd2589df |
| model-00017-of-00027.safetensors | Weights | 4.3 GB | b29732762a5f |
| model-00018-of-00027.safetensors | Weights | 4.3 GB | 187de89b0683 |
| model-00019-of-00027.safetensors | Weights | 4.4 GB | 8be152ac80c4 |
| model-00020-of-00027.safetensors | Weights | 4.3 GB | f8b7fb58cef5 |
| model-00021-of-00027.safetensors | Weights | 4.3 GB | 16d64cf1da41 |
| model-00022-of-00027.safetensors | Weights | 4.4 GB | d253517a5653 |
| model-00023-of-00027.safetensors | Weights | 4.2 GB | 0c819ea02c52 |
| model-00024-of-00027.safetensors | Weights | 4.1 GB | b641d71046ce |
| model-00025-of-00027.safetensors | Weights | 3.4 GB | 1ced1618186e |
| model-00026-of-00027.safetensors | Weights | 2.8 GB | 2fdda1da2d92 |
| model-00027-of-00027.safetensors | Weights | 1.9 GB | be3cb2424546 |
| model-vision-graft.safetensors | Weights | 912.1 MB | b47e83150554 |
| mtp-head-q6.safetensors | Weights | 5.8 GB | 1563d297c7bc |
| config.json | Configuration | 146.2 KB | — |
| generation_config.json | Configuration | 244 B | — |
| model.py | Configuration | 262.3 KB | — |
| model.safetensors.index.json | Configuration | 288.6 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| skipzero_load.py | Configuration | 5.6 KB | — |
| video_preprocessor_config.json | Configuration | 385 B | — |
| vqlab_provenance.json | Configuration | 8.1 KB | — |
| README.md | Documentation | 15.9 KB | — |
| chat_template.jinja | Other | 7.8 KB | — |
| mtp-head-q6.safetensors.fp32norms-bak | Other | 5.8 GB | a5cad76f094e |
| vqlab_provenance.history.jsonl | Other | 55.8 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 20.0 MB | 06b9509352d2 |
| tokenizer_config.json | Tokenizer | 1.2 KB | — |

## License and Download

License

apache-2.0

Access

Open weights, no gate

Download size

109.4 GB

[Download from Noah Zelezny](https://huggingface.co/TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.4bpw)

Released by Noah Zelezny through its official repository on Hugging Face. [Read the license](https://www.apache.org/licenses/LICENSE-2.0).

## Built From

- Derived from Qwen/Qwen3.5-397B-A17B
- Quantized from Qwen/Qwen3.5-397B-A17B

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 109.4 GB |
| 16-bit | 240.4 GB |
| 8-bit | 120.2 GB |
| 4-bit | 60.1 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About Qwen3.5-397B-A17B-VQ-2.4bpw

### How much GPU memory does Qwen3.5-397B-A17B-VQ-2.4bpw need?

About 288.4 GB at 16-bit and 72.1 GB at 4-bit: the weights (120.2B parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run Qwen3.5-397B-A17B-VQ-2.4bpw on?

At 16-bit, 2x MI300X from $3.70 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### Can I use Qwen3.5-397B-A17B-VQ-2.4bpw commercially?

Yes. Qwen3.5-397B-A17B-VQ-2.4bpw is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

### What is Qwen3.5-397B-A17B-VQ-2.4bpw's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

## Similar Models

Model · Text generation

### [keys-MiMo-V2.6-Pro-RL-Jarrelscy-ARVQ-Abliterated](https://savrn.com/models/keys-mimo-v2-6-pro-rl-jarrelscy-arvq-abliterated)

[Keys](https://savrn.com/model-publishers/drowzeys)

Abliterated Jarrelscy ARVQ / NVFP4 hybrid of XiaomiMiMo/MiMo-V2.6-Pro-RL. Thinking on/off is a request flag. Same weights. You choose per call. Thinking-off is the 100% gate. Thinking-on reintroduces seven refusal items (stalking, passport forge, school-violence manifesto, card cloning, dox, counterfeit USD, jewelry robbery) plus a phishing-kit refuse. Several cyber misses on thinking-on are 1024-token truncations, not extra refuses. Harmless probes stay clean in both modes. This model has had safety refusals removed. Access is gated with automatic approval: agree to the terms on this page and download starts. See RESPONSIBLEUSE.md. Xiaomi's chat template already supports both. Do not swap…

Access requested at publisher mit 119B parameters vllm

[View model](https://savrn.com/models/keys-mimo-v2-6-pro-rl-jarrelscy-arvq-abliterated)

Model · Text generation

### [gpt-oss-120b](https://savrn.com/models/gpt-oss-120b)

[OpenAI](https://savrn.com/model-publishers/openai)

Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. We’re releasing two flavors of these open models: - gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters) - gpt-oss-20b — for lower latency, and local or specialized use cases (21B parameters with 3.6B active parameters) Both models were trained on our harmony response format and should only be used with the harmony format as it will not work correctly otherwise. You can use gpt-oss-120b and gpt-oss-20b with Transformers.…

Open weights apache-2.0 116.8B parameters 131,072 tokens transformers

[View model](https://savrn.com/models/gpt-oss-120b)

Model · Text generation

### [NVIDIA-Nemotron-3-Super-120B-A12B-BF16](https://savrn.com/models/nvidia-nemotron-3-super-120b-a12b-bf16)

[NVIDIA](https://savrn.com/model-publishers/nvidia)

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-BF16 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Open weights other 123.6B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/nvidia-nemotron-3-super-120b-a12b-bf16)

Model · Text generation

### [Qwen3.8-Flash-Next-125B-A5B-INT4-AutoRound](https://savrn.com/models/qwen3-8-flash-next-125b-a5b-int4-autoround)

[Aldo Zampatti](https://savrn.com/model-publishers/azampatti)

Qwen3.8-Flash-Next with 5 routed experts per token instead of 10, healed so it stays close to the original, quantized to int4. It runs on one DGX Spark (GB10, 128 GB) at roughly 64-70 tokens/s. 125B parameters in total, 4.8B active per token. The original activates 6B. Everything needed to serve it is in this one repository, including the 49 GB FP8 n-gram table under ple-table/. Nothing else to download. That builds the serving image, downloads this repository, and starts an OpenAI-compatible server on port 8000. The scripts and the full explanation are in that repo. Serving by hand needs Saren-Arterius/qwen3.8-Flash-DGX-AutoRound, because a stock vLLM cannot serve this checkpoint's int4 +…

Open weights other 124B parameters 262,144 tokens vllm

[View model](https://savrn.com/models/qwen3-8-flash-next-125b-a5b-int4-autoround)

Model · Text generation

### [Vinci-Cyber-123B-1.0](https://savrn.com/models/vinci-cyber-123b-1-0)

[SimpleDirect](https://savrn.com/model-publishers/simpledirect)

Vinci Cyber 123B 1.0 is an open-weight model for defensive infrastructure review and targeted remediation, fine-tuned in Canada from Mistral AI's Devstral 2 123B. The released merged weights have now been tested directly, alongside their parent and the available GGUF formats. Focused repairs. Restraint on correct configuration. Weights you can run yourself. On the V2-B neutral-review test, the released BF16 model preserved 24/24 correct configurations and produced 18/24 scanner-credited repairs, all 18 passing offline provider-schema validation. Its parent repaired 17/24 and preserved 0/24. On the second set, V2-A, Cyber again preserved 24/24, but repaired 9/24 versus the parent's 15/24.…

Open weights other 125B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/vinci-cyber-123b-1-0)

Model · Text generation

### [Qwen3.8-Flash-Next-P48NVFP4-MoESQ](https://savrn.com/models/qwen3-8-flash-next-p48nvfp4-moesq)

[IST Austria Distributed Algorithms and Systems Lab](https://savrn.com/model-publishers/ista-daslab)

A W4A4 + paired-4:8 sparse compressed checkpoint of Qwen/Qwen3.8-Flash-Next, produced with MoESQ. The routed MoE expert weights are NVFP4 with paired-4:8 structured sparsity and are stored sparse: only the kept values plus a small mask are on disk. The target is NVIDIA Blackwell (SM100) sparse tensor cores. expert, 48 layers: 36 linear-attention and 12 sparse-attention, hyper-connections, per-layer n-gram embedding) weight including the sparsity mask and scales (2 bits per weight for the kept values). TP2). The per-layer n-gram embedding tables (95.4 GiB, BF16, unchanged) are lookup tables served from CPU memory, as in the base model. paired48nvfp4 MoE backend Links This checkpoint does not…

Open weights other 134.7B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/qwen3-8-flash-next-p48nvfp4-moesq)

## Noah Zelezny

[All models and datasets](https://savrn.com/model-publishers/thedrainflorist)

## Versions

- [34136c9eaf83](https://savrn.com/models/qwen3-5-397b-a17b-vq-2-4bpw/versions/34136c9eaf83) · current 2026-10-02

## Explore More

- [All text generation models](https://savrn.com/models/tasks/text-generation)
- [All models under apache-2.0](https://savrn.com/models/licenses/apache-2-0)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-10-02.
- [Hugging Face record](https://huggingface.co/TheDrainFlorist/Qwen3.5-397B-A17B-VQ-2.4bpw)
- [How the hub is built](https://savrn.com/model-hub/methodology)
