SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3.5-27B

by Qwen Qwen/Qwen3.5-27B

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance.

Parameters27.8B
Context262,144
Weights55.6 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1.9M

Runs On

What it takes to serve Qwen3.5-27B (27.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 55.6 GB 66.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 27.8 GB 33.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 13.9 GB 16.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

By Qwen, published under apache-2.0, revision fc05daec18b0.

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.

These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency.

Qwen3.5 Highlights

Qwen3.5 features the following enhancement:

Read the full model card (3,407 words)

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
64
Hidden size
5,120
Feed-forward size
17,408
Attention heads
24
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
Qwen/Qwen3.5-27B
Publisher
Qwen
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
27.8B parameters
Languages
Not stated by the source
Revision
fc05daec18b0a78c049392ed2e771dde82bdf654
First published
2026-02-24
Last updated
2026-04-24

Files and Weights

24 files, 55.6 GB in total. The weights are 11 files totalling 55.6 GB in safetensors.

Weights11 files · 55.6 GB
Configuration5 files · 131.8 KB
Tokenizer4 files · 22.9 MB
Documentation2 files · 103.8 KB
Other1 file · 7.8 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensors-00001-of-00011.safetensorsWeights5.3 GB 9019228d172c
model.safetensors-00002-of-00011.safetensorsWeights5.3 GB 890ef00c920b
model.safetensors-00003-of-00011.safetensorsWeights5.3 GB 8aca03689ad0
model.safetensors-00004-of-00011.safetensorsWeights5.3 GB 20f539430c60
model.safetensors-00005-of-00011.safetensorsWeights5.3 GB 57a0c074c654
model.safetensors-00006-of-00011.safetensorsWeights5.3 GB cfa4e6fbfc60
model.safetensors-00007-of-00011.safetensorsWeights5.3 GB d426963325b2
model.safetensors-00008-of-00011.safetensorsWeights5.4 GB fe60dbb9d253
model.safetensors-00009-of-00011.safetensorsWeights5.3 GB 71a153d88224
model.safetensors-00010-of-00011.safetensorsWeights5.3 GB 146745698b9f
model.safetensors-00011-of-00011.safetensorsWeights2.1 GB d947ce7483c4
config.jsonConfiguration4.1 KB
generation_config.jsonConfiguration244 B
model.safetensors.index.jsonConfiguration126.6 KB
preprocessor_config.jsonConfiguration390 B
video_preprocessor_config.jsonConfiguration385 B
LICENSEDocumentation11.3 KB
README.mdDocumentation92.4 KB
chat_template.jinjaOther7.8 KB
.gitattributesRepository1.6 KB
merges.txtTokenizer3.4 MB
tokenizer.jsonTokenizer12.8 MB 5f9e4d4901a9
tokenizer_config.jsonTokenizer16.7 KB
vocab.jsonTokenizer6.7 MB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
55.6 GB
Download from Qwen

Released by Qwen through ModelScope. Read the license.

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondSetup GPQA DiamondComparison conditions not established 81.8182 EvalEval
Reported by a third party
Evaluated revision not stated 2026-04-20
Idavidrein/gpqa Task diamondMetric diamondComparison conditions not established 85.5 Model Card
Reported by a third party
Evaluated revision not stated 2026-02-25
MMMU/MMMU_Pro Task mmmu_pro_visionMetric mmmu_pro_visionComparison conditions not established 75 Model Card
Reported by a third party
Evaluated revision not stated 2026-04-28
MathArena/aime_2026 Task MathArena/aime_2026Metric MathArena/aime_2026Comparison conditions not established 90.83 Official MathArena Evaluation
Reported by a third party
Evaluated revision not stated 2026-03-17
MathArena/hmmt_feb_2026 Task MathArena/hmmt_feb_2026Metric MathArena/hmmt_feb_2026Comparison conditions not established 81.06 Official MathArena Evaluation
Reported by a third party
Evaluated revision not stated 2026-03-17
SWE-bench/SWE-bench_Verified Task swe_bench_%_resolvedMetric swe_bench_%_resolvedComparison conditions not established 72.4 Model Card
Reported by a third party
Evaluated revision not stated 2026-02-03
TIGER-Lab/MMLU-Pro Task mmlu_proMetric mmlu_proComparison conditions not established 86.1 EvalEval
Reported by a third party
Evaluated revision not stated 2026-06-30
cais/hle Task hleMetric hleSetup with toolsComparison conditions not established 48.5 Model Card
Reported by a third party
Evaluated revision not stated 2026-02-25
cais/hle Task hleMetric hleSetup chain of thoughtComparison conditions not established 24.3 Model Card
Reported by a third party
Evaluated revision not stated 2026-02-25
harborframework/terminal-bench-2.0 Task terminalbench_2Metric terminalbench_2Comparison conditions not established 41.6 Model Card
Reported by a third party
Evaluated revision not stated 2026-02-25
likaixin/ScreenSpot-Pro Task overallMetric overallComparison conditions not established 70.3 Model Card
Reported by a third party
Evaluated revision not stated 2026-03-18

Memory Requirements

PrecisionWeights in memory
As published55.6 GB
16-bit55.6 GB
8-bit27.8 GB
4-bit13.9 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Hosted Prices

HostInput / outputUnitObserved
DeepInfra$0.26 / $2.60input / output, per million tokensSep 18, 2026
Novita$0.30 / $2.40input / output, per million tokensSep 18, 2026

From the SAVRN Index.

Compare Qwen3.5-27B

Questions About Qwen3.5-27B

How much GPU memory does Qwen3.5-27B need?

About 66.7 GB at 16-bit and 16.7 GB at 4-bit: the weights (27.8B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.5-27B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3.5-27B commercially?

Yes. Qwen3.5-27B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Qwen3.5-27B's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.8-27B

Qwen

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Qwen3.8 can be…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-FP8

Qwen

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Qwen3.8 can be…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.6-27B

Qwen

Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. This release delivers substantial upgrades, particularly in For more details, please refer to our blog post Qwen3.6-27B. Empty cells (--) indicate scores not yet available or not applicable. For streamlined integration, we recommend using Qwen3.6 via APIs. Below is a guide to use Qwen3.6 via OpenAI-compatible API. Qwen3.6 can be served via APIs with popular inference frameworks. In…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-AWQ-INT4

Cyankiwi

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Qwen3.8 can be…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B

Kyle Thomas

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Qwen3.8 can be…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Signal-3.8-27B-FP8

Vwdubb

This is Qwen3.8-27B that gets to the answer faster. AgentionAI Signal is a minimally invasive fine-tune of Qwen3.8-27B designed for lower generation latency and better token efficiency. On our held-out general-prompt evaluation, Signal produces 57% fewer answer tokens and uses 52% fewer thinking tokens, while matching or improving the measured answer quality of the base model. The percentages above were measured on the first release. The weights updated on 2026-09-13 trade a little of that reduction for stability; their re-measurement on the same prompt set is in progress and will replace these numbers. The result is substantially faster end-to-end generation: on typical chat prompts…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers