SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3.6-35B-A3B-NVFP4

by Red Hat AI RedHatAI/Qwen3.6-35B-A3B-NVFP4

NVFP4 variant of Qwen 3.6 35B for reasoning and tool calling.

Parameters34.7B
Context262,144
Weights25.0 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads967k

Runs On

What it takes to serve Qwen3.6-35B-A3B-NVFP4 (34.7B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 69.3 GB 83.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
8-bit 34.7 GB 41.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 17.3 GB 20.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on Qwen3.6-35B-A3B-NVFP4

Red Hat AI cut this one down for reasoning and tool calling, and the cut is the point. The full 34.7 billion parameters would need 83.2 GB at 16-bit; the 4-bit figure is 20.8 GB of memory on 17.3 GB of weights. Only 8 of the 256 experts fire per token, and the cheapest Index host, one MI300X with 192 GB at $1.85 an hour, leaves most of that card for context cache, which matters when the window runs to 262,144 tokens.

Apache 2.0 covers commercial use, modification and redistribution, with notices retained, significant changes stated and a patent grant included. Access is open. The relation to check is quantized_from Qwen/Qwen3.6-35B-A3B, updated August 13, 2026: an NVFP4 build inherits whatever the source model does, so test on your own prompts. The Index lists no per-token host price yet, so the cost case is GPU hours until one appears.

Model Card

By Red Hat AI, published under apache-2.0, revision e94a930c8e9b.

NVFP4 Quantized RedHatAI/Qwen3.6-35B-A3B-NVFP4

This is a preliminary version (and subject to change) of NVFP4 quantized Qwen/Qwen3.6-35B-A3B model. The model has both weights and activations quantized to NVFP4 format with vllm-project/llm-compressor.

It is compatible and tested against vllm main. Deploy it with: vllm serve RedHatAI/Qwen3.6-35B-A3B-NVFP4 --reasoning-parser qwen3 --moe_backend flashinfer_cutlass.

If you have hardware with more compute than memory bandwidth, you may prefer this MoE variant for performance reasons.

Creation Script:

Run this script with LLM Compressor main and latest transformers.

Read the full model card (690 words)

Configuration

Architecture
Qwen3_5MoeForConditionalGeneration
Context length (tokens)
262,144
Layers
40
Hidden size
2,048
Attention heads
16
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Experts
256
Experts active per token
8
Model type
qwen3_5_moe
Quantization
compressed-tensors

Identity and Version

Repository
RedHatAI/Qwen3.6-35B-A3B-NVFP4
Publisher
Red Hat AI
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
34.7B parameters
Languages
Not stated by the source
Revision
e94a930c8e9b8a7d3763e449289c261c9c7d3e5e
First published
2026-04-17
Last updated
2026-08-13

Files and Weights

87 files, 25.1 GB in total. The weights are 3 files totalling 25.0 GB in safetensors.

Weights3 files · 25.0 GB
Configuration65 files · 57.7 MB
Tokenizer2 files · 20.0 MB
Documentation1 file · 9.7 KB
Other15 files · 5.8 MB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights22.5 GB bb759388bf46
model_mtp.safetensorsWeights1.7 GB 68b2cabe2ff8
model_visual.safetensorsWeights893.2 MB 1e1b1d5f576d
base_score/Qwen_Qwen3.6-35B-A3B/agentic/BFCL_v4_web_search_base_score.jsonConfiguration1.8 MB
base_score/Qwen_Qwen3.6-35B-A3B/agentic/BFCL_v4_web_search_no_snippet_score.jsonConfiguration3.2 MB
base_score/Qwen_Qwen3.6-35B-A3B/agentic/memory/kv/BFCL_v4_memory_kv_score.jsonConfiguration857.8 KB
base_score/Qwen_Qwen3.6-35B-A3B/agentic/memory/rec_sum/BFCL_v4_memory_rec_sum_score.jsonConfiguration504.9 KB
base_score/Qwen_Qwen3.6-35B-A3B/agentic/memory/vector/BFCL_v4_memory_vector_score.jsonConfiguration1.1 MB
base_score/Qwen_Qwen3.6-35B-A3B/live/BFCL_v4_live_irrelevance_score.jsonConfiguration567.4 KB
base_score/Qwen_Qwen3.6-35B-A3B/live/BFCL_v4_live_multiple_score.jsonConfiguration1.3 MB
base_score/Qwen_Qwen3.6-35B-A3B/live/BFCL_v4_live_parallel_multiple_score.jsonConfiguration46.7 KB
base_score/Qwen_Qwen3.6-35B-A3B/live/BFCL_v4_live_parallel_score.jsonConfiguration10.9 KB
base_score/Qwen_Qwen3.6-35B-A3B/live/BFCL_v4_live_relevance_score.jsonConfiguration5.7 KB
base_score/Qwen_Qwen3.6-35B-A3B/live/BFCL_v4_live_simple_score.jsonConfiguration185.9 KB
base_score/Qwen_Qwen3.6-35B-A3B/multi_turn/BFCL_v4_multi_turn_base_score.jsonConfiguration929.7 KB
base_score/Qwen_Qwen3.6-35B-A3B/multi_turn/BFCL_v4_multi_turn_long_context_score.jsonConfiguration5.5 MB
base_score/Qwen_Qwen3.6-35B-A3B/multi_turn/BFCL_v4_multi_turn_miss_func_score.jsonConfiguration1.7 MB
base_score/Qwen_Qwen3.6-35B-A3B/multi_turn/BFCL_v4_multi_turn_miss_param_score.jsonConfiguration2.0 MB
base_score/Qwen_Qwen3.6-35B-A3B/non_live/BFCL_v4_irrelevance_score.jsonConfiguration26.4 KB
base_score/Qwen_Qwen3.6-35B-A3B/non_live/BFCL_v4_multiple_score.jsonConfiguration278.4 KB
base_score/Qwen_Qwen3.6-35B-A3B/non_live/BFCL_v4_parallel_multiple_score.jsonConfiguration669.6 KB
base_score/Qwen_Qwen3.6-35B-A3B/non_live/BFCL_v4_parallel_score.jsonConfiguration270.7 KB
base_score/Qwen_Qwen3.6-35B-A3B/non_live/BFCL_v4_simple_java_score.jsonConfiguration155.0 KB
base_score/Qwen_Qwen3.6-35B-A3B/non_live/BFCL_v4_simple_javascript_score.jsonConfiguration25.4 KB
base_score/Qwen_Qwen3.6-35B-A3B/non_live/BFCL_v4_simple_python_score.jsonConfiguration246.1 KB
bf16_every_eval_ever/aime25.jsonConfiguration3.8 KB
bf16_every_eval_ever/gpqa_diamond.jsonConfiguration2.4 KB
bf16_every_eval_ever/gsm8k_platinum_cot_llama.jsonConfiguration3.9 KB
bf16_every_eval_ever/ifeval.jsonConfiguration6.2 KB
bf16_every_eval_ever/lcb_codegeneration_v6.jsonConfiguration2.4 KB
bf16_every_eval_ever/math_500.jsonConfiguration2.2 KB
bf16_every_eval_ever/mmlu_pro_chat.jsonConfiguration1.8 KB
bf16_every_eval_ever/swebench.jsonConfiguration43.1 KB
config.jsonConfiguration23.2 KB
generation_config.jsonConfiguration213 B
model.safetensors.index.jsonConfiguration12.4 MB 3a43cf79383a
nvfp4_every_eval_ever/aime25.jsonConfiguration3.8 KB
nvfp4_every_eval_ever/gpqa_diamond.jsonConfiguration2.4 KB
nvfp4_every_eval_ever/gsm8k_platinum_cot_llama.jsonConfiguration4.0 KB
nvfp4_every_eval_ever/ifeval.jsonConfiguration6.2 KB
nvfp4_every_eval_ever/lcb_codegeneration_v6.jsonConfiguration2.4 KB
nvfp4_every_eval_ever/math_500.jsonConfiguration2.2 KB
nvfp4_every_eval_ever/mmlu_pro_chat.jsonConfiguration1.8 KB
nvfp4_every_eval_ever/swebench.jsonConfiguration42.2 KB
processor_config.jsonConfiguration1.2 KB
recipe.yamlConfiguration311 B
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/agentic/BFCL_v4_web_search_base_score.jsonConfiguration2.9 MB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/agentic/BFCL_v4_web_search_no_snippet_score.jsonConfiguration3.3 MB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/agentic/memory/kv/BFCL_v4_memory_kv_score.jsonConfiguration1.0 MB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/agentic/memory/rec_sum/BFCL_v4_memory_rec_sum_score.jsonConfiguration567.1 KB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/agentic/memory/vector/BFCL_v4_memory_vector_score.jsonConfiguration986.6 KB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/live/BFCL_v4_live_irrelevance_score.jsonConfiguration582.1 KB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/live/BFCL_v4_live_multiple_score.jsonConfiguration1.3 MB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/live/BFCL_v4_live_parallel_multiple_score.jsonConfiguration54.1 KB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/live/BFCL_v4_live_parallel_score.jsonConfiguration13.4 KB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/live/BFCL_v4_live_relevance_score.jsonConfiguration2.0 KB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/live/BFCL_v4_live_simple_score.jsonConfiguration207.3 KB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/multi_turn/BFCL_v4_multi_turn_base_score.jsonConfiguration1.1 MB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/multi_turn/BFCL_v4_multi_turn_long_context_score.jsonConfiguration5.7 MB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/multi_turn/BFCL_v4_multi_turn_miss_func_score.jsonConfiguration2.1 MB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/multi_turn/BFCL_v4_multi_turn_miss_param_score.jsonConfiguration2.1 MB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/non_live/BFCL_v4_irrelevance_score.jsonConfiguration36.2 KB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/non_live/BFCL_v4_multiple_score.jsonConfiguration280.9 KB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/non_live/BFCL_v4_parallel_multiple_score.jsonConfiguration671.8 KB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/non_live/BFCL_v4_parallel_score.jsonConfiguration268.7 KB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/non_live/BFCL_v4_simple_java_score.jsonConfiguration155.1 KB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/non_live/BFCL_v4_simple_javascript_score.jsonConfiguration21.0 KB
score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/non_live/BFCL_v4_simple_python_score.jsonConfiguration241.1 KB
README.mdDocumentation9.7 KB
base_score/data_agentic.csvOther241 B
base_score/data_format_sensitivity.csvOther2.7 KB
base_score/data_live.csvOther252 B
base_score/data_multi_turn.csvOther135 B
base_score/data_non_live.csvOther277 B
base_score/data_overall.csvOther966 B
chat_template.jinjaOther7.8 KB
results_35B-A3B_NVFP4_all.txtOther2.9 MB
results_35B-A3B_base_all.txtOther2.9 MB
score/data_agentic.csvOther251 B
score/data_format_sensitivity.csvOther2.7 KB
score/data_live.csvOther262 B
score/data_multi_turn.csvOther145 B
score/data_non_live.csvOther287 B
score/data_overall.csvOther986 B
.gitattributesRepository1.6 KB
tokenizer.jsonTokenizer20.0 MB dd6b8cf757c2
tokenizer_config.jsonTokenizer1.1 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
25.0 GB
Download from Red Hat AI

Released by Red Hat AI through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published25.0 GB
16-bit69.3 GB
8-bit34.7 GB
4-bit17.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare Qwen3.6-35B-A3B-NVFP4

Questions About Qwen3.6-35B-A3B-NVFP4

How much GPU memory does Qwen3.6-35B-A3B-NVFP4 need?

About 83.2 GB at 16-bit and 20.8 GB at 4-bit: the weights (34.7B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.6-35B-A3B-NVFP4 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3.6-35B-A3B-NVFP4 commercially?

Yes. Qwen3.6-35B-A3B-NVFP4 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Qwen3.6-35B-A3B-NVFP4's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen2.5-VL-32B-Instruct-AWQ

Qwen

In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to better align with human preferences. Particularly for objective queries such as mathematics, logical reasoning, and knowledge-based Q&A, the level of detail in responses and the clarity of formatting have been noticeably enhanced. In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on…

Open weights apache-2.0 33.5B parameters 128,000 tokens transformers

Model · Image and text to text

Qwen2.5-VL-32B-Instruct

Qwen

In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to better align with human preferences. Particularly for objective queries such as mathematics, logical reasoning, and knowledge-based Q&A, the level of detail in responses and the clarity of formatting have been noticeably enhanced. In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on…

Open weights apache-2.0 33.5B parameters 128,000 tokens transformers

Model · Image and text to text

Qwen3.6-35B-A3B

Qwen

Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. This release delivers substantial upgrades, particularly in For more details, please refer to our blog post Qwen3.6-35B-A3B. Empty cells (--) indicate scores not available or not applicable. For streamlined integration, we recommend using Qwen3.6 via APIs. Below is a guide to use Qwen3.6 via OpenAI-compatible API. Qwen3.6 can be served via APIs with popular inference frameworks. In…

Open weights apache-2.0 36B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.5-35B-A3B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 36B parameters 262,144 tokens transformers

Base: Qwen/Qwen3.6-35B-A3B-FP8 (same FP8 block format/scales convention). Adapter: qwen3.6-35b-a3b-tool-prompts-ckpt444-adapter (LoRA r=32, 2 epochs, 1873 dialogues). Merged in BF16 (RAM only), block-requantized to FP8 e4m3/128. Chat template: qwen3template333.jinja. Serve with sglang / vllm / transformers.

Open weights apache-2.0 36B parameters 262,144 tokens