SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Qwen3.5-397B-A17B-VQ-3.1bpw

by Noah Zelezny TheDrainFlorist/Qwen3.5-397B-A17B-VQ-3.1bpw

Qwen3.5-397B-A17B-VQ-3.1bpw is an open-weight model for text generation from Noah Zelezny, released under Apache License 2.0. It has 64.5B parameters and a 262,144-token context. At 16-bit it needs about 154.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 1.7k downloads a month.

125.3 GiB text weights — the quality build. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 131.6 GiB.

Parameters64.5B
Context262,144
Weights141.3 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1.7k

Runs On

What it takes to serve Qwen3.5-397B-A17B-VQ-3.1bpw (64.5B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 128.9 GB 154.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
8-bit 64.5 GB 77.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 32.2 GB 38.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Qwen3.5-397B-A17B-VQ-3.1bpw on every accelerator the SAVRN Index prices, at every precision

Model Card

By Noah Zelezny, published under apache-2.0, revision 54a606613d27.

125.3 GiB text weights — the quality build. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 131.6 GiB.

A vector-quantized build of Qwen3.5-397B-A17B for machines with memory to spend: the strongest quantization we know how to make of this model at this size, on stock mlx-lm, no patches.

Changelog

2026-10-01 — vq-skipzero: dead rows dropped

12.43% of the 64-element weight groups in the Qwen3.5-397B teacher sit in output rows whose weights are about 1e-29. The vq-skipzero format drops those rows' codes and scales on disk and in memory; they output exact zeros, and the live rows are byte-identical to the previous revision (vqlab sz-check, every module). Text weights go from 141.71 to 125.31 GiB (-16.40 GiB), and resident memory drops by about the same amount. This revision has not been KL-scored; the live rows are byte-identical (vqlab sz-check) and generation was verified.

Read the full model card (2,921 words)

Configuration

Architecture
Qwen3_5MoeForConditionalGeneration
Context length (tokens)
262,144
Layers
60
Hidden size
4,096
Attention heads
32
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Experts
512
Experts active per token
10
Model type
qwen3_5_moe

Identity and Version

Repository
TheDrainFlorist/Qwen3.5-397B-A17B-VQ-3.1bpw
Publisher
Noah Zelezny
Task
Text generation
Modality
Text
Library
mlx
Parameters
64.5B parameters
Languages
en
Revision
54a606613d27d1163f5e05ea67740014cb162c05
First published
2026-08-16
Last updated
2026-10-02

Files and Weights

44 files, 147.1 GB in total. The weights are 29 files totalling 141.3 GB in safetensors.

Weights29 files · 141.3 GB
Configuration8 files · 715.6 KB
Tokenizer2 files · 20.0 MB
Documentation1 file · 19.8 KB
Other3 files · 5.8 GB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00027.safetensorsWeights3.2 GB 387a5e64f1d5
model-00002-of-00027.safetensorsWeights3.4 GB bf605a5e381f
model-00003-of-00027.safetensorsWeights4.2 GB a30fc662fcc4
model-00004-of-00027.safetensorsWeights4.7 GB fc47d6a61428
model-00005-of-00027.safetensorsWeights4.5 GB 54035c82fb0a
model-00006-of-00027.safetensorsWeights4.6 GB 97a92733b02e
model-00007-of-00027.safetensorsWeights5.1 GB e5931fc67b9d
model-00008-of-00027.safetensorsWeights5.0 GB 19cd4a199193
model-00009-of-00027.safetensorsWeights5.1 GB 5cf99241bfb9
model-00010-of-00027.safetensorsWeights5.5 GB 1f1edc036d59
model-00011-of-00027.safetensorsWeights5.4 GB c56d3fa81da8
model-00012-of-00027.safetensorsWeights5.5 GB 09a0b802fc53
model-00013-of-00027.safetensorsWeights5.7 GB 6c1adda21f9b
model-00014-of-00027.safetensorsWeights5.7 GB 2d3da44a5bf7
model-00015-of-00027.safetensorsWeights5.7 GB 65967a2b087b
model-00016-of-00027.safetensorsWeights5.8 GB 0809bf57a62e
model-00017-of-00027.safetensorsWeights5.7 GB 917dc3c50f0f
model-00018-of-00027.safetensorsWeights5.6 GB db60c728ee80
model-00019-of-00027.safetensorsWeights5.8 GB 6ad8198a2ac0
model-00020-of-00027.safetensorsWeights5.7 GB 008b331d0620
model-00021-of-00027.safetensorsWeights5.6 GB 8a8fa571a98b
model-00022-of-00027.safetensorsWeights5.7 GB 8fff7efee5f8
model-00023-of-00027.safetensorsWeights5.5 GB 358c0c9ba128
model-00024-of-00027.safetensorsWeights5.4 GB 52aef7420cd4
model-00025-of-00027.safetensorsWeights4.5 GB abb5603db56c
model-00026-of-00027.safetensorsWeights3.7 GB 71e10b5487a9
model-00027-of-00027.safetensorsWeights2.2 GB 311f6d45cac3
model-vision-graft.safetensorsWeights912.1 MB b47e83150554
mtp-head-q6.safetensorsWeights5.8 GB 1563d297c7bc
config.jsonConfiguration150.0 KB —
generation_config.jsonConfiguration244 B —
model.pyConfiguration262.3 KB —
model.safetensors.index.jsonConfiguration288.6 KB —
preprocessor_config.jsonConfiguration390 B —
skipzero_load.pyConfiguration5.6 KB —
video_preprocessor_config.jsonConfiguration385 B —
vqlab_provenance.jsonConfiguration8.1 KB —
README.mdDocumentation19.8 KB —
chat_template.jinjaOther7.8 KB —
mtp-head-q6.safetensors.fp32norms-bakOther5.8 GB a5cad76f094e
vqlab_provenance.history.jsonlOther57.3 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.2 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
141.3 GB
Download from Noah Zelezny

Released by Noah Zelezny through its official repository on Hugging Face. Read the license.

Built From

  • Derived from Qwen/Qwen3.5-397B-A17B
  • Quantized from Qwen/Qwen3.5-397B-A17B

Memory Requirements

PrecisionWeights in memory
As published141.3 GB
16-bit128.9 GB
8-bit64.5 GB
4-bit32.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3.5-397B-A17B-VQ-3.1bpw

How much GPU memory does Qwen3.5-397B-A17B-VQ-3.1bpw need?

About 154.7 GB at 16-bit and 38.7 GB at 4-bit: the weights (64.5B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.5-397B-A17B-VQ-3.1bpw on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3.5-397B-A17B-VQ-3.1bpw commercially?

Yes. Qwen3.5-397B-A17B-VQ-3.1bpw is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Qwen3.5-397B-A17B-VQ-3.1bpw's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

Qwen3.5-122B-A10B-NVFP4

NVIDIA

The NVIDIA Qwen3.5-122B-A10B-NVFP4 model is the quantized version of Alibaba's Qwen3.5-122B-A10B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.5-122B-A10B NVFP4 model is quantized with Model Optimizer. This model is ready for commercial/non-commercial use. This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case; see link to Non-NVIDIA (Qwen3.5-122B-A10B) Model Card from Alibaba. Global Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems…

Open weights apache-2.0 64.6B parameters 262,144 tokens Model Optimizer

This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,632 tokens. The source checkpoint was loaded and pruned in BF16. The source-row selection hash is…

Open weights apache-2.0 64.1B parameters 262,144 tokens

This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,632 tokens. The source checkpoint was loaded and pruned in BF16. The source-row selection hash is…

Open weights apache-2.0 64.1B parameters 262,144 tokens

This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 64 calibration samples from the gpqamain configuration of Idavidrein/gpqa revision 633f5ee89ab8ad4522a9f850766b73f62147ffdd. The released loader settings are preserved: train, Question plus shuffled choices, Explanation, selectionseed=1234, BF16, and no optimizer step. The samples are full length: no tokenizer maxlength, truncation, or padding. The longest input for this tokenizer is 1,632 tokens. The source checkpoint was loaded and pruned in BF16. The source-row selection hash is…

Open weights apache-2.0 64.1B parameters 262,144 tokens

Model · Text generation

Qwen3.5-397B-A17B-VQ-2.2bpw

Noah Zelezny

88.7 GiB text weights — the accessibility build, the roomiest fit on a 128 GB Mac. Also ships the bf16 vision tower (+0.85 GiB) and an optional MTP draft head (+5.4 GiB); full download 95.0 GiB. (v2, mixed geometry.) v2 — updated 2026-08-22. This repository now serves a rebuilt artifact at the same size and the same bits per weight, with a different codebook geometry that measures better on both perplexity corpora. v1's numbers are kept below rather than quietly overwritten, and v1's bytes remain downloadable by pinning the previous revision: A vector-quantized build of Qwen3.5-397B-A17B running? 88.7 GiB text weights — it runs on a single 128 GB Apple Silicon machine with ≈7 GiB more…

Open weights apache-2.0 62.6B parameters 262,144 tokens mlx

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-NVFP4 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Open weights other 67.2B parameters 262,144 tokens transformers