Open-weight model · Image and text to text
Qwen3.6-35B-A3B-NVFP4
by Red Hat AI RedHatAI/Qwen3.6-35B-A3B-NVFP4
NVFP4 variant of Qwen 3.6 35B for reasoning and tool calling.
Runs On
What it takes to serve Qwen3.6-35B-A3B-NVFP4 (34.7B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 69.3 GB | 83.2 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x MI325X $2.00 · 1x MI355X $2.59 |
| 8-bit | 34.7 GB | 41.6 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 17.3 GB | 20.8 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.
SAVRN's Notes on Qwen3.6-35B-A3B-NVFP4
Red Hat AI cut this one down for reasoning and tool calling, and the cut is the point. The full 34.7 billion parameters would need 83.2 GB at 16-bit; the 4-bit figure is 20.8 GB of memory on 17.3 GB of weights. Only 8 of the 256 experts fire per token, and the cheapest Index host, one MI300X with 192 GB at $1.85 an hour, leaves most of that card for context cache, which matters when the window runs to 262,144 tokens.
Apache 2.0 covers commercial use, modification and redistribution, with notices retained, significant changes stated and a patent grant included. Access is open. The relation to check is quantized_from Qwen/Qwen3.6-35B-A3B, updated August 13, 2026: an NVFP4 build inherits whatever the source model does, so test on your own prompts. The Index lists no per-token host price yet, so the cost case is GPU hours until one appears.
Model Card
By Red Hat AI, published under apache-2.0, revision e94a930c8e9b.
NVFP4 Quantized RedHatAI/Qwen3.6-35B-A3B-NVFP4
This is a preliminary version (and subject to change) of NVFP4 quantized Qwen/Qwen3.6-35B-A3B model. The model has both weights and activations quantized to NVFP4 format with vllm-project/llm-compressor.
It is compatible and tested against vllm main. Deploy it with: vllm serve RedHatAI/Qwen3.6-35B-A3B-NVFP4 --reasoning-parser qwen3 --moe_backend flashinfer_cutlass.
If you have hardware with more compute than memory bandwidth, you may prefer this MoE variant for performance reasons.
Creation Script:
Run this script with LLM Compressor main and latest transformers.
Configuration
- Architecture
- Qwen3_5MoeForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 40
- Hidden size
- 2,048
- Attention heads
- 16
- Key/value heads
- 2
- Head dimension
- 256
- Vocabulary size
- 248,320
- Experts
- 256
- Experts active per token
- 8
- Model type
- qwen3_5_moe
- Quantization
- compressed-tensors
Identity and Version
- Repository
- RedHatAI/Qwen3.6-35B-A3B-NVFP4
- Publisher
- Red Hat AI
- Task
- Image and text to text
- Modality
- Image and text
- Library
- transformers
- Parameters
- 34.7B parameters
- Languages
- Not stated by the source
- Revision
- e94a930c8e9b8a7d3763e449289c261c9c7d3e5e
- First published
- 2026-04-17
- Last updated
- 2026-08-13
Files and Weights
87 files, 25.1 GB in total. The weights are 3 files totalling 25.0 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 22.5 GB | bb759388bf46 |
| model_mtp.safetensors | Weights | 1.7 GB | 68b2cabe2ff8 |
| model_visual.safetensors | Weights | 893.2 MB | 1e1b1d5f576d |
| base_score/Qwen_Qwen3.6-35B-A3B/agentic/BFCL_v4_web_search_base_score.json | Configuration | 1.8 MB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/agentic/BFCL_v4_web_search_no_snippet_score.json | Configuration | 3.2 MB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/agentic/memory/kv/BFCL_v4_memory_kv_score.json | Configuration | 857.8 KB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/agentic/memory/rec_sum/BFCL_v4_memory_rec_sum_score.json | Configuration | 504.9 KB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/agentic/memory/vector/BFCL_v4_memory_vector_score.json | Configuration | 1.1 MB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/live/BFCL_v4_live_irrelevance_score.json | Configuration | 567.4 KB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/live/BFCL_v4_live_multiple_score.json | Configuration | 1.3 MB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/live/BFCL_v4_live_parallel_multiple_score.json | Configuration | 46.7 KB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/live/BFCL_v4_live_parallel_score.json | Configuration | 10.9 KB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/live/BFCL_v4_live_relevance_score.json | Configuration | 5.7 KB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/live/BFCL_v4_live_simple_score.json | Configuration | 185.9 KB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/multi_turn/BFCL_v4_multi_turn_base_score.json | Configuration | 929.7 KB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/multi_turn/BFCL_v4_multi_turn_long_context_score.json | Configuration | 5.5 MB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/multi_turn/BFCL_v4_multi_turn_miss_func_score.json | Configuration | 1.7 MB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/multi_turn/BFCL_v4_multi_turn_miss_param_score.json | Configuration | 2.0 MB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/non_live/BFCL_v4_irrelevance_score.json | Configuration | 26.4 KB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/non_live/BFCL_v4_multiple_score.json | Configuration | 278.4 KB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/non_live/BFCL_v4_parallel_multiple_score.json | Configuration | 669.6 KB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/non_live/BFCL_v4_parallel_score.json | Configuration | 270.7 KB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/non_live/BFCL_v4_simple_java_score.json | Configuration | 155.0 KB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/non_live/BFCL_v4_simple_javascript_score.json | Configuration | 25.4 KB | — |
| base_score/Qwen_Qwen3.6-35B-A3B/non_live/BFCL_v4_simple_python_score.json | Configuration | 246.1 KB | — |
| bf16_every_eval_ever/aime25.json | Configuration | 3.8 KB | — |
| bf16_every_eval_ever/gpqa_diamond.json | Configuration | 2.4 KB | — |
| bf16_every_eval_ever/gsm8k_platinum_cot_llama.json | Configuration | 3.9 KB | — |
| bf16_every_eval_ever/ifeval.json | Configuration | 6.2 KB | — |
| bf16_every_eval_ever/lcb_codegeneration_v6.json | Configuration | 2.4 KB | — |
| bf16_every_eval_ever/math_500.json | Configuration | 2.2 KB | — |
| bf16_every_eval_ever/mmlu_pro_chat.json | Configuration | 1.8 KB | — |
| bf16_every_eval_ever/swebench.json | Configuration | 43.1 KB | — |
| config.json | Configuration | 23.2 KB | — |
| generation_config.json | Configuration | 213 B | — |
| model.safetensors.index.json | Configuration | 12.4 MB | 3a43cf79383a |
| nvfp4_every_eval_ever/aime25.json | Configuration | 3.8 KB | — |
| nvfp4_every_eval_ever/gpqa_diamond.json | Configuration | 2.4 KB | — |
| nvfp4_every_eval_ever/gsm8k_platinum_cot_llama.json | Configuration | 4.0 KB | — |
| nvfp4_every_eval_ever/ifeval.json | Configuration | 6.2 KB | — |
| nvfp4_every_eval_ever/lcb_codegeneration_v6.json | Configuration | 2.4 KB | — |
| nvfp4_every_eval_ever/math_500.json | Configuration | 2.2 KB | — |
| nvfp4_every_eval_ever/mmlu_pro_chat.json | Configuration | 1.8 KB | — |
| nvfp4_every_eval_ever/swebench.json | Configuration | 42.2 KB | — |
| processor_config.json | Configuration | 1.2 KB | — |
| recipe.yaml | Configuration | 311 B | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/agentic/BFCL_v4_web_search_base_score.json | Configuration | 2.9 MB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/agentic/BFCL_v4_web_search_no_snippet_score.json | Configuration | 3.3 MB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/agentic/memory/kv/BFCL_v4_memory_kv_score.json | Configuration | 1.0 MB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/agentic/memory/rec_sum/BFCL_v4_memory_rec_sum_score.json | Configuration | 567.1 KB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/agentic/memory/vector/BFCL_v4_memory_vector_score.json | Configuration | 986.6 KB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/live/BFCL_v4_live_irrelevance_score.json | Configuration | 582.1 KB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/live/BFCL_v4_live_multiple_score.json | Configuration | 1.3 MB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/live/BFCL_v4_live_parallel_multiple_score.json | Configuration | 54.1 KB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/live/BFCL_v4_live_parallel_score.json | Configuration | 13.4 KB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/live/BFCL_v4_live_relevance_score.json | Configuration | 2.0 KB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/live/BFCL_v4_live_simple_score.json | Configuration | 207.3 KB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/multi_turn/BFCL_v4_multi_turn_base_score.json | Configuration | 1.1 MB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/multi_turn/BFCL_v4_multi_turn_long_context_score.json | Configuration | 5.7 MB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/multi_turn/BFCL_v4_multi_turn_miss_func_score.json | Configuration | 2.1 MB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/multi_turn/BFCL_v4_multi_turn_miss_param_score.json | Configuration | 2.1 MB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/non_live/BFCL_v4_irrelevance_score.json | Configuration | 36.2 KB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/non_live/BFCL_v4_multiple_score.json | Configuration | 280.9 KB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/non_live/BFCL_v4_parallel_multiple_score.json | Configuration | 671.8 KB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/non_live/BFCL_v4_parallel_score.json | Configuration | 268.7 KB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/non_live/BFCL_v4_simple_java_score.json | Configuration | 155.1 KB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/non_live/BFCL_v4_simple_javascript_score.json | Configuration | 21.0 KB | — |
| score/RedHatAI_Qwen3.6-35B-A3B-NVFP4/non_live/BFCL_v4_simple_python_score.json | Configuration | 241.1 KB | — |
| README.md | Documentation | 9.7 KB | — |
| base_score/data_agentic.csv | Other | 241 B | — |
| base_score/data_format_sensitivity.csv | Other | 2.7 KB | — |
| base_score/data_live.csv | Other | 252 B | — |
| base_score/data_multi_turn.csv | Other | 135 B | — |
| base_score/data_non_live.csv | Other | 277 B | — |
| base_score/data_overall.csv | Other | 966 B | — |
| chat_template.jinja | Other | 7.8 KB | — |
| results_35B-A3B_NVFP4_all.txt | Other | 2.9 MB | — |
| results_35B-A3B_base_all.txt | Other | 2.9 MB | — |
| score/data_agentic.csv | Other | 251 B | — |
| score/data_format_sensitivity.csv | Other | 2.7 KB | — |
| score/data_live.csv | Other | 262 B | — |
| score/data_multi_turn.csv | Other | 145 B | — |
| score/data_non_live.csv | Other | 287 B | — |
| score/data_overall.csv | Other | 986 B | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 20.0 MB | dd6b8cf757c2 |
| tokenizer_config.json | Tokenizer | 1.1 KB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 25.0 GB
Released by Red Hat AI through its official repository on Hugging Face. Read the license.
Built From
- Derived from Qwen/Qwen3.6-35B-A3B
- Quantized from Qwen/Qwen3.6-35B-A3B
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 25.0 GB |
| 16-bit | 69.3 GB |
| 8-bit | 34.7 GB |
| 4-bit | 17.3 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Compare Qwen3.6-35B-A3B-NVFP4
Questions About Qwen3.6-35B-A3B-NVFP4
How much GPU memory does Qwen3.6-35B-A3B-NVFP4 need?
About 83.2 GB at 16-bit and 20.8 GB at 4-bit: the weights (34.7B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Qwen3.6-35B-A3B-NVFP4 on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use Qwen3.6-35B-A3B-NVFP4 commercially?
Yes. Qwen3.6-35B-A3B-NVFP4 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is Qwen3.6-35B-A3B-NVFP4's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.
Similar Models
In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to better align with human preferences. Particularly for objective queries such as mathematics, logical reasoning, and knowledge-based Q&A, the level of detail in responses and the clarity of formatting have been noticeably enhanced. In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on…
In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to better align with human preferences. Particularly for objective queries such as mathematics, logical reasoning, and knowledge-based Q&A, the level of detail in responses and the clarity of formatting have been noticeably enhanced. In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on…
Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. This release delivers substantial upgrades, particularly in For more details, please refer to our blog post Qwen3.6-35B-A3B. Empty cells (--) indicate scores not available or not applicable. For streamlined integration, we recommend using Qwen3.6 via APIs. Below is a guide to use Qwen3.6 via OpenAI-compatible API. Qwen3.6 can be served via APIs with popular inference frameworks. In…
Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…
Base: Qwen/Qwen3.6-35B-A3B-FP8 (same FP8 block format/scales convention). Adapter: qwen3.6-35b-a3b-tool-prompts-ckpt444-adapter (LoRA r=32, 2 epochs, 1873 dialogues). Merged in BF16 (RAM only), block-requantized to FP8 e4m3/128. Chat template: qwen3template333.jinja. Serve with sglang / vllm / transformers.