SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen2.5-VL-32B-Instruct

by Qwen Qwen/Qwen2.5-VL-32B-Instruct

In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning.

Parameters33.5B
Context128,000
Weights68.3 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1.1M

Runs On

What it takes to serve Qwen2.5-VL-32B-Instruct (33.5B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 66.9 GB 80.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
8-bit 33.5 GB 40.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 16.7 GB 20.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on Qwen2.5-VL-32B-Instruct

Run it at 16-bit and you need 80.3 GB, so precision is the hardware decision. At 8-bit it needs 40.1 GB, at 4-bit 20.1 GB. The cheapest setup on our board is one 192 GB MI300X at $1.85 per hour on-demand, and at 16-bit the weights alone take 66.9 GB of that card before context loads. Images and text go in, text comes out. Qwen ships the 33.5B parameters as 32 safetensors files, about 68.3 GB, so plan the storage pull first.

Apache 2.0 covers commercial use, modification and redistribution, provided notices stay and significant changes are stated, which suits a deployment inside your facility. Check the context: 128,000 tokens is the ceiling, the configuration lists a 32,768 token sliding window, and the page cites the YaRN paper on context window extension, arXiv 2309.00071, so test with your longest real inputs before committing. The technical report is arXiv 2502.13923.

Model Card

By Qwen, published under apache-2.0, revision 7cfb30d71a1f.

Latest Updates:

In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to better align with human preferences. Particularly for objective queries such as mathematics, logical reasoning, and knowledge-based Q&A, the level of detail in responses and the clarity of formatting have been noticeably enhanced.

Introduction

In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introduce the latest addition to the Qwen family: Qwen2.5-VL.

Key Enhancements:

Read the full model card (2,138 words)

Configuration

Architecture
Qwen2_5_VLForConditionalGeneration
Context length (tokens)
128,000
Layers
64
Hidden size
5,120
Feed-forward size
27,648
Attention heads
40
Key/value heads
8
Vocabulary size
152,064
Sliding window (tokens)
32,768
RoPE base
1e+06
Stored precision
bfloat16
Model type
qwen2_5_vl

Identity and Version

Repository
Qwen/Qwen2.5-VL-32B-Instruct
Publisher
Qwen
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
33.5B parameters
Languages
en
Revision
7cfb30d71a1f4f49a57592323337a4a4727301da
First published
2025-03-21
Last updated
2025-04-14

Files and Weights

32 files, 68.3 GB in total. The weights are 18 files totalling 68.3 GB in safetensors.

Weights18 files · 68.3 GB
Configuration8 files · 97.2 KB
Tokenizer4 files · 15.9 MB
Documentation1 file · 18.6 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00018.safetensorsWeights2.8 GB 64c2a870b88b
model-00002-of-00018.safetensorsWeights3.9 GB d58b007f2a2a
model-00003-of-00018.safetensorsWeights3.9 GB e4102551d612
model-00004-of-00018.safetensorsWeights3.9 GB 20a6ff2c6b6e
model-00005-of-00018.safetensorsWeights3.9 GB 3af4c2e3274e
model-00006-of-00018.safetensorsWeights3.9 GB 8a47aa90e2e6
model-00007-of-00018.safetensorsWeights3.9 GB dcd94eb28bee
model-00008-of-00018.safetensorsWeights3.9 GB 5ed4aee864b4
model-00009-of-00018.safetensorsWeights3.9 GB bd7a7666f724
model-00010-of-00018.safetensorsWeights3.9 GB 23e22a974e69
model-00011-of-00018.safetensorsWeights3.9 GB 2d9900a158e7
model-00012-of-00018.safetensorsWeights3.9 GB 7d0005e8ea7c
model-00013-of-00018.safetensorsWeights3.9 GB 792a18eb6999
model-00014-of-00018.safetensorsWeights3.9 GB d3ef548f5c10
model-00015-of-00018.safetensorsWeights3.9 GB f99201514697
model-00016-of-00018.safetensorsWeights3.9 GB e366d738ecd9
model-00017-of-00018.safetensorsWeights3.9 GB 92e355aac688
model-00018-of-00018.safetensorsWeights3.1 GB 87d7dddef2c0
added_tokens.jsonConfiguration605 B
chat_template.jsonConfiguration1.0 KB
config.jsonConfiguration1.2 KB
configuration.jsonConfiguration76 B
generation_config.jsonConfiguration216 B
model.safetensors.index.jsonConfiguration93.1 KB
preprocessor_config.jsonConfiguration351 B
special_tokens_map.jsonConfiguration613 B
README.mdDocumentation18.6 KB
.gitattributesRepository1.6 KB
merges.txtTokenizer1.7 MB
tokenizer.jsonTokenizer11.4 MB 9c5ae00e602b
tokenizer_config.jsonTokenizer5.7 KB
vocab.jsonTokenizer2.8 MB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
68.3 GB
Download from Qwen

Released by Qwen through ModelScope. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published68.3 GB
16-bit66.9 GB
8-bit33.5 GB
4-bit16.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Compare Qwen2.5-VL-32B-Instruct

Questions About Qwen2.5-VL-32B-Instruct

How much GPU memory does Qwen2.5-VL-32B-Instruct need?

About 80.3 GB at 16-bit and 20.1 GB at 4-bit: the weights (33.5B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen2.5-VL-32B-Instruct on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen2.5-VL-32B-Instruct commercially?

Yes. Qwen2.5-VL-32B-Instruct is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Qwen2.5-VL-32B-Instruct's context length?

128,000 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen2.5-VL-32B-Instruct-AWQ

Qwen

In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to better align with human preferences. Particularly for objective queries such as mathematics, logical reasoning, and knowledge-based Q&A, the level of detail in responses and the clarity of formatting have been noticeably enhanced. In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on…

Open weights apache-2.0 33.5B parameters 128,000 tokens transformers

Model · Image and text to text

gemma-4-31B-it

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0 31.3B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.6-35B-A3B

Qwen

Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. This release delivers substantial upgrades, particularly in For more details, please refer to our blog post Qwen3.6-35B-A3B. Empty cells (--) indicate scores not available or not applicable. For streamlined integration, we recommend using Qwen3.6 via APIs. Below is a guide to use Qwen3.6 via OpenAI-compatible API. Qwen3.6 can be served via APIs with popular inference frameworks. In…

Open weights apache-2.0 36B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.5-35B-A3B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 36B parameters 262,144 tokens transformers