SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3.8-27B

by Qwen Qwen/Qwen3.8-27B

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date.

Parameters27.8B
Context262,144
Weights55.6 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads7.4M

Runs On

What it takes to serve Qwen3.8-27B (27.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 55.6 GB 66.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 27.8 GB 33.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 13.9 GB 16.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on Qwen3.8-27B

Plan around 66.7 GB for Qwen3.8-27B at 16-bit. That fits one MI300X with 192 GB, which the SAVRN Index prices at $1.85 an hour on demand, and 8-bit brings it to 33.3 GB, 4-bit to 16.7 GB, all on that one card. It reads images and text and writes text, with a 262,144-token context, so it belongs where one accelerator per instance is the budget and the inputs include pictures or long documents.

Apache 2.0 allows commercial use, modification and redistribution, provided the license and notice files stay with the weights and significant changes are stated. Weigh the card against renting tokens on the Index: Cerebras at $0.99 in and $1.49 out per million, DeepInfra at $0.20 and $2.50, Novita at $0.42 and $3.00, OVHcloud at $0.47 and $3.19. The benchmark figures on the page are reported by the model card and outside evaluators, not measured by SAVRN.

Model Card

By Qwen, published under apache-2.0, revision 1d4bf0f2ff60.

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.

These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc.

[!Tip] For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. In particular, Qwen3.8-27B will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates.

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date.

Read the full model card (2,690 words)

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
64
Hidden size
5,120
Feed-forward size
17,408
Attention heads
24
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
Qwen/Qwen3.8-27B
Publisher
Qwen
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
27.8B parameters
Languages
Not stated by the source
Revision
1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
First published
2026-08-05
Last updated
2026-08-14

Files and Weights

32 files, 55.6 GB in total. The weights are 18 files totalling 55.6 GB in safetensors.

Weights18 files · 55.6 GB
Configuration5 files · 117.5 KB
Tokenizer4 files · 22.9 MB
Documentation2 files · 76.6 KB
Other2 files · 9.2 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00018.safetensorsWeights4.0 GB ba0ce20aae48
model-00002-of-00018.safetensorsWeights3.0 GB 06a148c01bfb
model-00003-of-00018.safetensorsWeights2.5 GB 2e1bf62cbcd4
model-00004-of-00018.safetensorsWeights4.0 GB 511e34063187
model-00005-of-00018.safetensorsWeights2.1 GB 635cb53446dc
model-00006-of-00018.safetensorsWeights4.0 GB 0bc5214fac60
model-00007-of-00018.safetensorsWeights2.1 GB 80b0c49033e9
model-00008-of-00018.safetensorsWeights4.0 GB 7192c5b66185
model-00009-of-00018.safetensorsWeights2.1 GB af3c48cc37af
model-00010-of-00018.safetensorsWeights4.0 GB 163490a76f3b
model-00011-of-00018.safetensorsWeights2.1 GB 5f3ae1b948ae
model-00012-of-00018.safetensorsWeights4.0 GB a3de1c711467
model-00013-of-00018.safetensorsWeights2.1 GB 06ab79a41f74
model-00014-of-00018.safetensorsWeights4.0 GB 4138ed946030
model-00015-of-00018.safetensorsWeights2.1 GB 69224e27b9de
model-00016-of-00018.safetensorsWeights4.0 GB 73cb9a1089fb
model-00017-of-00018.safetensorsWeights2.1 GB beb51f010561
model-00018-of-00018.safetensorsWeights3.4 GB 1d3479509e21
config.jsonConfiguration4.3 KB
generation_config.jsonConfiguration202 B
model.safetensors.index.jsonConfiguration112.2 KB
preprocessor_config.jsonConfiguration390 B
video_preprocessor_config.jsonConfiguration385 B
LICENSEDocumentation11.5 KB
README.mdDocumentation65.0 KB
chat_template.jinjaOther9.0 KB
crc32.txtOther238 B
.gitattributesRepository1.6 KB
merges.txtTokenizer3.4 MB
tokenizer.jsonTokenizer12.8 MB 0997f410c57a
tokenizer_config.jsonTokenizer17.9 KB
vocab.jsonTokenizer6.7 MB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
55.6 GB
Download from Qwen

Released by Qwen through ModelScope. Read the license.

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Idavidrein/gpqa Task diamondMetric diamondComparison conditions not established 89.2 Qwen3.8-27B model card
Reported by a third party
Evaluated revision not stated 2026-08-14
ScaleAI/SWE-bench_Pro Task SWE_Bench_ProMetric SWE_Bench_ProSetup Evaluated with the Claude Code harness, temp=1.0, top_p=0.95, 256K context; baseline models re-evaluated on the same refined task set.Comparison conditions not established 61.7 Qwen3.8-27B model card
Reported by a third party
Evaluated revision not stated 2026-08-14
cais/hle Task hleMetric hleSetup Judged by GPT-4o.Comparison conditions not established 30.8 Qwen3.8-27B model card
Reported by a third party
Evaluated revision not stated 2026-08-14
claw-eval/Claw-Eval Task multimodalMetric multimodalSetup Reported as ClawEval-MM. Pass@3, the benchmark's own Pass³ methodology; card also reports a secondary 56.9 'Average' metric, not included here.Comparison conditions not established 57.4 Qwen3.8-27B model card
Reported by a third party
Evaluated revision not stated 2026-08-14
datacurve/deep-swe Task deep_sweMetric deep_sweComparison conditions not established 42.2 Qwen3.8-27B model card
Reported by a third party
Evaluated revision not stated 2026-08-14
harborframework/terminal-bench-2.1 Task terminalbench_2_1Metric terminalbench_2_1Setup Row labeled "(Terminus)" as the harness; no further hyperparameter footnote given for this row.Comparison conditions not established 73 Model Card
Reported by a third party
Evaluated revision not stated 2026-08-14
internlm/WildClawBench Task avg_timeMetric avg_timeComparison conditions not established 516 WildClawBench
Reported by a third party
Evaluated revision not stated 2026-08-16
internlm/WildClawBench Task overallMetric overallComparison conditions not established 48.0152 WildClawBench
Reported by a third party
Evaluated revision not stated 2026-08-16
llamaindex/ExtractBench Task longMetric longSetup Pipeline name: qwen3_8_27b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established 38.45 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-24
llamaindex/ExtractBench Task meanMetric meanSetup Pipeline name: qwen3_8_27b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established 89.75 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-24
llamaindex/ExtractBench Task mediumMetric mediumSetup Pipeline name: qwen3_8_27b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established 87.54 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-24
llamaindex/ExtractBench Task shortMetric shortSetup Pipeline name: qwen3_8_27b_fp8_vllm_extract_oneshot_structured_output_file; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established 94.68 ExtractBench
Reported by a third party
Evaluated revision not stated 2026-08-24
llamaindex/ParseBench Task chartMetric chartSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established 69.17 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-28
llamaindex/ParseBench Task layoutMetric layoutSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established 69.9 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-28
llamaindex/ParseBench Task meanMetric meanSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established 70.79 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-28
llamaindex/ParseBench Task tableMetric tableSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established 66.82 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-28
llamaindex/ParseBench Task text_contentMetric text_contentSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established 88.28 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-28
llamaindex/ParseBench Task text_formattingMetric text_formattingSetup Pipeline name: qwen3_8_27b_thinking_parse_with_layout; served checkpoint: Qwen/Qwen3.8-27B-FP8Comparison conditions not established 59.77 ParseBench
Reported by a third party
Evaluated revision not stated 2026-08-28

Memory Requirements

PrecisionWeights in memory
As published55.6 GB
16-bit55.6 GB
8-bit27.8 GB
4-bit13.9 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Hosted Prices

HostInput / outputUnitObserved
Cerebras$0.99 / $1.49input / output, per million tokensSep 18, 2026
DeepInfra$0.20 / $2.50input / output, per million tokensSep 18, 2026
Novita$0.42 / $3.00input / output, per million tokensSep 18, 2026
OVHcloud$0.47 / $3.19input / output, per million tokensSep 18, 2026

From the SAVRN Index.

Built on This Model

Compare Qwen3.8-27B

Questions About Qwen3.8-27B

How much GPU memory does Qwen3.8-27B need?

About 66.7 GB at 16-bit and 16.7 GB at 4-bit: the weights (27.8B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.8-27B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen3.8-27B commercially?

Yes. Qwen3.8-27B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Qwen3.8-27B's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.8-27B-FP8

Qwen

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Qwen3.8 can be…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.6-27B

Qwen

Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. This release delivers substantial upgrades, particularly in For more details, please refer to our blog post Qwen3.6-27B. Empty cells (--) indicate scores not yet available or not applicable. For streamlined integration, we recommend using Qwen3.6 via APIs. Below is a guide to use Qwen3.6 via OpenAI-compatible API. Qwen3.6 can be served via APIs with popular inference frameworks. In…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.5-27B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-AWQ-INT4

Cyankiwi

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Qwen3.8 can be…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B

Kyle Thomas

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Qwen3.8 can be…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Signal-3.8-27B-FP8

Vwdubb

This is Qwen3.8-27B that gets to the answer faster. AgentionAI Signal is a minimally invasive fine-tune of Qwen3.8-27B designed for lower generation latency and better token efficiency. On our held-out general-prompt evaluation, Signal produces 57% fewer answer tokens and uses 52% fewer thinking tokens, while matching or improving the measured answer quality of the base model. The percentages above were measured on the first release. The weights updated on 2026-09-13 trade a little of that reduction for stability; their re-measurement on the same prompt set is in progress and will replace these numbers. The result is substantially faster end-to-end generation: on typical chat prompts…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers