SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3.5-35B-A3B

by Qwen Qwen/Qwen3.5-35B-A3B

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance.

Parameters36B
Context262,144
Weights71.9 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1.9M

Runs On

What it takes to serve Qwen3.5-35B-A3B (36B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 71.9 GB 86.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
8-bit 36.0 GB 43.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 18.0 GB 21.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on Qwen3.5-35B-A3B

Two hosts on the SAVRN Index sell this model by the token, DeepInfra at $0.14 in and $1.00 out per million and Novita at $0.25 and $2.00, so rent and own compare directly. Images and text in, text out, across a 262,144-token context and 256 experts, 8 active per token. At 16-bit it needs 86.3 GB for 71.9 GB of weights, fitting one MI300X with 192 GB at $1.85 per hour on demand; 43.1 GB at 8-bit and 21.6 GB at 4-bit leave most of the card for context.

Under Apache 2.0 you can modify, redistribute and sell what you build, notices kept, changes stated. It is derived from Qwen3.5-35B-A3B-Base, so decide whether post-training starts here or at the base. Then price the 262,144-token window, where a 192 GB card's headroom goes, and set your hourly cost against the two Index prices at your real volume before you buy.

Model Card

By Qwen, published under apache-2.0, revision 59d61f3ce65a.

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.

These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

[!Tip] For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Alibaba Cloud Model Studio.

In particular, Qwen3.5-Flash is the hosted version corresponding to Qwen3.5-35B-A3B with more production features, e.g., 1M context length by default and official built-in tools. For more information, please refer to the User Guide.

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency.

Qwen3.5 Highlights

Qwen3.5 features the following enhancement:

Read the full model card (3,483 words)

Configuration

Architecture
Qwen3_5MoeForConditionalGeneration
Context length (tokens)
262,144
Layers
40
Hidden size
2,048
Attention heads
16
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Experts
256
Experts active per token
8
Model type
qwen3_5_moe

Identity and Version

Repository
Qwen/Qwen3.5-35B-A3B
Publisher
Qwen
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
36B parameters
Languages
Not stated by the source
Revision
59d61f3ce65a6d9863b86d2e96597125219dc754
First published
2026-02-24
Last updated
2026-04-24

Files and Weights

27 files, 71.9 GB in total. The weights are 14 files totalling 71.9 GB in safetensors.

Weights14 files · 71.9 GB
Configuration5 files · 192.0 KB
Tokenizer4 files · 22.9 MB
Documentation2 files · 104.7 KB
Other1 file · 7.8 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensors-00001-of-00014.safetensorsWeights5.4 GB 1b9f16fdae24
model.safetensors-00002-of-00014.safetensorsWeights5.4 GB 79f06bc069d2
model.safetensors-00003-of-00014.safetensorsWeights5.4 GB 2cb56343806f
model.safetensors-00004-of-00014.safetensorsWeights5.4 GB 39a39c799e52
model.safetensors-00005-of-00014.safetensorsWeights5.4 GB 1a3b31ebb04b
model.safetensors-00006-of-00014.safetensorsWeights5.4 GB af5f17d292df
model.safetensors-00007-of-00014.safetensorsWeights5.4 GB 52885dbecd4e
model.safetensors-00008-of-00014.safetensorsWeights5.4 GB 43f232ba029d
model.safetensors-00009-of-00014.safetensorsWeights5.3 GB 8fa6a375dd53
model.safetensors-00010-of-00014.safetensorsWeights5.4 GB d89e80aaf84c
model.safetensors-00011-of-00014.safetensorsWeights5.4 GB 5b774025e5ed
model.safetensors-00012-of-00014.safetensorsWeights5.4 GB d0f0bd7a5f09
model.safetensors-00013-of-00014.safetensorsWeights5.4 GB da8cc27a2a99
model.safetensors-00014-of-00014.safetensorsWeights2.2 GB d5e08a7dd670
config.jsonConfiguration3.5 KB
generation_config.jsonConfiguration244 B
model.safetensors.index.jsonConfiguration187.5 KB
preprocessor_config.jsonConfiguration390 B
video_preprocessor_config.jsonConfiguration385 B
LICENSEDocumentation11.5 KB
README.mdDocumentation93.2 KB
chat_template.jinjaOther7.8 KB
.gitattributesRepository1.6 KB
merges.txtTokenizer3.4 MB
tokenizer.jsonTokenizer12.8 MB 5f9e4d4901a9
tokenizer_config.jsonTokenizer16.7 KB
vocab.jsonTokenizer6.7 MB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
71.9 GB
Download from Qwen

Released by Qwen through ModelScope. Read the license.

Built From

  • Derived from Qwen/Qwen3.5-35B-A3B-Base

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondSetup GPQA DiamondComparison conditions not established 85.3535 EvalEval
Reported by a third party
Evaluated revision not stated 2026-04-20
Idavidrein/gpqa Task diamondMetric diamondComparison conditions not established 84.2 Model Card
Reported by a third party
Evaluated revision not stated 2026-02-25
MMMU/MMMU_Pro Task mmmu_pro_visionMetric mmmu_pro_visionComparison conditions not established 75.1 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-28
MathArena/aime_2026 Task MathArena/aime_2026Metric MathArena/aime_2026Comparison conditions not established 93.33 Official MathArena Evaluation
Reported by a third party
Evaluated revision not stated 2026-03-17
MathArena/hmmt_feb_2026 Task MathArena/hmmt_feb_2026Metric MathArena/hmmt_feb_2026Comparison conditions not established 81.82 Official MathArena Evaluation
Reported by a third party
Evaluated revision not stated 2026-03-17
SWE-bench/SWE-bench_Verified Task swe_bench_%_resolvedMetric swe_bench_%_resolvedComparison conditions not established 69.2 Model Card
Reported by a third party
Evaluated revision not stated 2026-02-03
TIGER-Lab/MMLU-Pro Task mmlu_proMetric mmlu_proComparison conditions not established 85.3 EvalEval
Reported by a third party
Evaluated revision not stated 2026-06-30
cais/hle Task hleMetric hleSetup chain of thoughtComparison conditions not established 22.4 Model Card
Reported by a third party
Evaluated revision not stated 2026-02-25
harborframework/terminal-bench-2.0 Task terminalbench_2Metric terminalbench_2Comparison conditions not established 40.5 Model Card
Reported by a third party
Evaluated revision not stated 2026-02-25
joelniklaus/LEXam-hard Task lexam_hardMetric lexam_hardSetup lighteval, LEXam paper prompts, one response per question, no tools; DeepSeek-R1-0528 judge; mean of the German and English means over the 518 questions, 0-100Comparison conditions not established 23.93 SwissLegalEvals per-sample details (lighteval)
Reported by a third party
Evaluated revision not stated 2026-08-01
likaixin/ScreenSpot-Pro Task overallMetric overallComparison conditions not established 68.6 Model Card
Reported by a third party
Evaluated revision not stated 2026-03-18
llamaindex/ExtractBench Task longMetric longSetup Pipeline name: qwen3_5_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.5-35B-A3B-FP8Comparison conditions not established 31.73 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-24
llamaindex/ExtractBench Task meanMetric meanSetup Pipeline name: qwen3_5_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.5-35B-A3B-FP8Comparison conditions not established 88 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-24
llamaindex/ExtractBench Task mediumMetric mediumSetup Pipeline name: qwen3_5_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.5-35B-A3B-FP8Comparison conditions not established 85.27 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-24
llamaindex/ExtractBench Task shortMetric shortSetup Pipeline name: qwen3_5_35b_a3b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.5-35B-A3B-FP8Comparison conditions not established 93.52 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-24

Memory Requirements

PrecisionWeights in memory
As published71.9 GB
16-bit71.9 GB
8-bit36.0 GB
4-bit18.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Hosted Prices

HostInput / outputUnitObserved
DeepInfra$0.14 / $1.00input / output, per million tokensSep 18, 2026
Novita$0.25 / $2.00input / output, per million tokensSep 18, 2026

From the SAVRN Index.

Built on This Model

Compare Qwen3.5-35B-A3B

Questions About Qwen3.5-35B-A3B

How much GPU memory does Qwen3.5-35B-A3B need?

About 86.3 GB at 16-bit and 21.6 GB at 4-bit: the weights (36B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.5-35B-A3B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3.5-35B-A3B commercially?

Yes. Qwen3.5-35B-A3B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Qwen3.5-35B-A3B's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.6-35B-A3B

Qwen

Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. This release delivers substantial upgrades, particularly in For more details, please refer to our blog post Qwen3.6-35B-A3B. Empty cells (--) indicate scores not available or not applicable. For streamlined integration, we recommend using Qwen3.6 via APIs. Below is a guide to use Qwen3.6 via OpenAI-compatible API. Qwen3.6 can be served via APIs with popular inference frameworks. In…

Open weights apache-2.0 36B parameters 262,144 tokens transformers

Base: Qwen/Qwen3.6-35B-A3B-FP8 (same FP8 block format/scales convention). Adapter: qwen3.6-35b-a3b-tool-prompts-ckpt444-adapter (LoRA r=32, 2 epochs, 1873 dialogues). Merged in BF16 (RAM only), block-requantized to FP8 e4m3/128. Chat template: qwen3template333.jinja. Serve with sglang / vllm / transformers.

Open weights apache-2.0 36B parameters 262,144 tokens

Model · Image and text to text

Qwen3.6-35B-A3B-FP8

Qwen

Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. This release delivers substantial upgrades, particularly in For more details, please refer to our blog post Qwen3.6-35B-A3B. Empty cells (--) indicate scores not available or not applicable. For streamlined integration, we recommend using Qwen3.6 via APIs. Below is a guide to use Qwen3.6 via OpenAI-compatible API. Qwen3.6 can be served via APIs with popular inference frameworks. In…

Open weights apache-2.0 36B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.5-35B-A3B-FP8

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 36B parameters 262,144 tokens transformers