SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3.6-35B-A3B

by Qwen Qwen/Qwen3.6-35B-A3B

Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6.

Parameters36B
Context262,144
Weights71.9 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads3.4M

Runs On

What it takes to serve Qwen3.6-35B-A3B (36B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 71.9 GB 86.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
8-bit 36.0 GB 43.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 18.0 GB 21.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on Qwen3.6-35B-A3B

86.3 GB is the figure that picks the hardware. The 16-bit weights need that much once loaded, and the cheapest listed card, one MI300X with 192 GB, holds it at $1.85 an hour on-demand. At 4-bit the footprint is 21.6 GB, leaving most of the card for the 262,144-token window. Only 8 of 256 experts fire per token: memory for all 36B parameters, compute for a slice. It takes images as well as text, and the publisher's summary leads with coding.

Apache 2.0 allows commercial use, modification and redistribution, provided the notices travel with it. Before committing a card, price the rental: DeepInfra lists $0.10 in and $0.95 out per million tokens, Scaleway $0.29 and $1.71. If your volume will not fill $1.85 an hour, rent; if it will, the 71.9 GB of safetensors is yours. Released April 15, 2026 and updated April 24, so check the revision you pulled.

Model Card

By Qwen, published under apache-2.0, revision 995ad96eacd9.

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.

These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience.

Qwen3.6 Highlights

This release delivers substantial upgrades, particularly in

  • Agentic Coding: the model now handles frontend workflows and repository-level reasoning with greater fluency and precision.
  • Thinking Preservation: we've introduced a new option to retain reasoning context from historical messages, streamlining iterative development and reducing overhead.

For more details, please refer to our blog post Qwen3.6-35B-A3B.

Model Overview

Read the full model card (3,106 words)

Configuration

Architecture
Qwen3_5MoeForConditionalGeneration
Context length (tokens)
262,144
Layers
40
Hidden size
2,048
Attention heads
16
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Experts
256
Experts active per token
8
Model type
qwen3_5_moe

Identity and Version

Repository
Qwen/Qwen3.6-35B-A3B
Publisher
Qwen
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
36B parameters
Languages
Not stated by the source
Revision
995ad96eacd98c81ed38be0c5b274b04031597b0
First published
2026-04-15
Last updated
2026-04-24

Files and Weights

40 files, 71.9 GB in total. The weights are 26 files totalling 71.9 GB in safetensors.

Weights26 files · 71.9 GB
Configuration6 files · 103.1 KB
Tokenizer4 files · 22.9 MB
Documentation2 files · 75.9 KB
Other1 file · 7.8 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00026.safetensorsWeights4.0 GB adee7bcb930a
model-00002-of-00026.safetensorsWeights1.3 GB 88f2dfd2b9e7
model-00003-of-00026.safetensorsWeights3.4 GB 8f7d72178d3f
model-00004-of-00026.safetensorsWeights3.4 GB 12d7db38689b
model-00005-of-00026.safetensorsWeights3.4 GB a836047305d0
model-00006-of-00026.safetensorsWeights4.0 GB c9080d718e9c
model-00007-of-00026.safetensorsWeights1.1 GB e8c05e23131b
model-00008-of-00026.safetensorsWeights3.9 GB 4b6a6d495053
model-00009-of-00026.safetensorsWeights1.1 GB a31a954bb72d
model-00010-of-00026.safetensorsWeights3.9 GB 246560e66570
model-00011-of-00026.safetensorsWeights1.1 GB 7180392817fe
model-00012-of-00026.safetensorsWeights3.4 GB 043fb525f662
model-00013-of-00026.safetensorsWeights1.6 GB 33a20fb20a21
model-00014-of-00026.safetensorsWeights3.4 GB be823e33c5cb
model-00015-of-00026.safetensorsWeights1.6 GB a89d547c6f9d
model-00016-of-00026.safetensorsWeights3.9 GB 69fc3ae03164
model-00017-of-00026.safetensorsWeights1.1 GB e356e3943cf3
model-00018-of-00026.safetensorsWeights3.9 GB 9e5e63fd1cc7
model-00019-of-00026.safetensorsWeights1.1 GB 708644ad34f1
model-00020-of-00026.safetensorsWeights3.4 GB ca083a1d1aa6
model-00021-of-00026.safetensorsWeights1.6 GB ada4ae48f3d4
model-00022-of-00026.safetensorsWeights3.4 GB def207fb42d7
model-00023-of-00026.safetensorsWeights3.4 GB 864d52ca7768
model-00024-of-00026.safetensorsWeights3.4 GB 391acd27420c
model-00025-of-00026.safetensorsWeights3.8 GB 778e7f76602f
model-00026-of-00026.safetensorsWeights2.2 GB 1a9740422007
config.jsonConfiguration3.7 KB
configuration.jsonConfiguration58 B
generation_config.jsonConfiguration202 B
model.safetensors.index.jsonConfiguration98.4 KB
preprocessor_config.jsonConfiguration390 B
video_preprocessor_config.jsonConfiguration385 B
LICENSEDocumentation11.3 KB
README.mdDocumentation64.5 KB
chat_template.jinjaOther7.8 KB
.gitattributesRepository1.6 KB
merges.txtTokenizer3.4 MB
tokenizer.jsonTokenizer12.8 MB 5f9e4d4901a9
tokenizer_config.jsonTokenizer16.7 KB
vocab.jsonTokenizer6.7 MB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
71.9 GB
Download from Qwen

Released by Qwen through ModelScope. Read the license.

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondComparison conditions not established 86 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-16
MMMU/MMMU_Pro Task mmmu_pro_visionMetric mmmu_pro_visionComparison conditions not established 75.3 Model Card
Reported by a third party
Evaluated revision not stated 2026-05-15
MathArena/aime_2026 Task MathArena/aime_2026Metric MathArena/aime_2026Comparison conditions not established 92.7 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-16
MathArena/hmmt_feb_2026 Task MathArena/hmmt_feb_2026Metric MathArena/hmmt_feb_2026Comparison conditions not established 83.6 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-16
SWE-bench/SWE-bench_Multilingual Task swe_bench_multilingual_%_resolvedMetric swe_bench_multilingual_%_resolvedComparison conditions not established 67.2 Model Card
Reported by a third party
Evaluated revision not stated 2026-08-10
SWE-bench/SWE-bench_Verified Task swe_bench_%_resolvedMetric swe_bench_%_resolvedComparison conditions not established 73.4 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-16
ScaleAI/SWE-bench_Pro Task SWE_Bench_ProMetric SWE_Bench_ProComparison conditions not established 49.5 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-16
TIGER-Lab/MMLU-Pro Task mmlu_proMetric mmlu_proComparison conditions not established 85.2 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-16
cais/hle Task hleMetric hleComparison conditions not established 21.4 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-16
harborframework/terminal-bench-2.0 Task terminalbench_2Metric terminalbench_2Setup Harbor/Terminus-2 harness; 3h timeout, 32 CPU/48 GB RAM; temp=1.0, top_p=0.95, top_k=20, max_tokens=80K, 256K ctx; avg of 5 runs.Comparison conditions not established 51.5 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-16
llamaindex/ExtractBench Task longMetric longSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint Qwen/Qwen3.6-35B-A3B-FP8 on vLLM, one-shot json_object structured outputComparison conditions not established 36.9 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ExtractBench Task meanMetric meanSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint Qwen/Qwen3.6-35B-A3B-FP8 on vLLM, one-shot json_object structured outputComparison conditions not established 88.11 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ExtractBench Task mediumMetric mediumSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint Qwen/Qwen3.6-35B-A3B-FP8 on vLLM, one-shot json_object structured outputComparison conditions not established 86.37 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ExtractBench Task shortMetric shortSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint Qwen/Qwen3.6-35B-A3B-FP8 on vLLM, one-shot json_object structured outputComparison conditions not established 92.86 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-26
llamaindex/ParseBench Task chartMetric chartSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_parse_layoutComparison conditions not established 5.1 ParseBench
Reported by a third party
Evaluated revision not stated 2026-04-16
llamaindex/ParseBench Task layoutMetric layoutSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_parse_layoutComparison conditions not established 47.4 ParseBench
Reported by a third party
Evaluated revision not stated 2026-04-16
llamaindex/ParseBench Task meanMetric meanSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_parse_layoutComparison conditions not established 44.1 ParseBench
Reported by a third party
Evaluated revision not stated 2026-04-16
llamaindex/ParseBench Task tableMetric tableSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_parse_layoutComparison conditions not established 19.1 ParseBench
Reported by a third party
Evaluated revision not stated 2026-04-16
llamaindex/ParseBench Task text_contentMetric text_contentSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_parse_layoutComparison conditions not established 90.7 ParseBench
Reported by a third party
Evaluated revision not stated 2026-04-16
llamaindex/ParseBench Task text_formattingMetric text_formattingSetup Pipeline name: qwen3_6_35b_a3b_fp8_vllm_parse_layoutComparison conditions not established 58.3 ParseBench
Reported by a third party
Evaluated revision not stated 2026-04-16

Memory Requirements

PrecisionWeights in memory
As published71.9 GB
16-bit71.9 GB
8-bit36.0 GB
4-bit18.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Hosted Prices

HostInput / outputUnitObserved
DeepInfra$0.10 / $0.95input / output, per million tokensSep 18, 2026
Scaleway$0.29 / $1.71input / output, per million tokensSep 18, 2026

From the SAVRN Index.

Built on This Model

Compare Qwen3.6-35B-A3B

Questions About Qwen3.6-35B-A3B

How much GPU memory does Qwen3.6-35B-A3B need?

About 86.3 GB at 16-bit and 21.6 GB at 4-bit: the weights (36B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.6-35B-A3B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3.6-35B-A3B commercially?

Yes. Qwen3.6-35B-A3B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Qwen3.6-35B-A3B's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.5-35B-A3B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 36B parameters 262,144 tokens transformers

Base: Qwen/Qwen3.6-35B-A3B-FP8 (same FP8 block format/scales convention). Adapter: qwen3.6-35b-a3b-tool-prompts-ckpt444-adapter (LoRA r=32, 2 epochs, 1873 dialogues). Merged in BF16 (RAM only), block-requantized to FP8 e4m3/128. Chat template: qwen3template333.jinja. Serve with sglang / vllm / transformers.

Open weights apache-2.0 36B parameters 262,144 tokens

Model · Image and text to text

Qwen3.6-35B-A3B-FP8

Qwen

Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. This release delivers substantial upgrades, particularly in For more details, please refer to our blog post Qwen3.6-35B-A3B. Empty cells (--) indicate scores not available or not applicable. For streamlined integration, we recommend using Qwen3.6 via APIs. Below is a guide to use Qwen3.6 via OpenAI-compatible API. Qwen3.6 can be served via APIs with popular inference frameworks. In…

Open weights apache-2.0 36B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.5-35B-A3B-FP8

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 36B parameters 262,144 tokens transformers