SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen2-VL-7B-Instruct

by Qwen Qwen/Qwen2-VL-7B-Instruct

We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation.

Parameters8.3B
Context32,768
Weights16.6 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads931.7k

Runs On

What it takes to serve Qwen2-VL-7B-Instruct (8.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 16.6 GB 19.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 8.3 GB 9.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 4.1 GB 5.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on Qwen2-VL-7B-Instruct

Nineteen point nine gigabytes at 16-bit. That figure decides where Qwen2-VL-7B-Instruct lands, and it fits on one MI300X with room left over, at $1.85 an hour, the cheapest rate we track. At 8-bit the footprint drops to 9.9 GB and at 4-bit to 5.0 GB. The work is images and text in, text out: Qwen built it to read pictures at varied resolutions and aspect ratios and to answer questions about video longer than 20 minutes, with a 32,768-token context to hold what it sees.

Apache 2.0 allows commercial use, modification and redistribution, with notices kept. Before committing, note that the instruct weights derive from the Qwen2-VL-7B base, that the only benchmark in our record is a third-party ScreenSpot-Pro result of 1.6 overall, and that no host on the SAVRN Index prices it per token, so the hourly figure above is the cost basis.

Model Card

By Qwen, published under apache-2.0, revision eed13092ef92.

Introduction

We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation.

What’s New in Qwen2-VL?

Key Enhancements:
  • SoTA understanding of images of various resolution & ratio: Qwen2-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc.

  • Understanding videos of 20min+: Qwen2-VL can understand videos over 20 minutes for high-quality video-based question answering, dialog, content creation, etc.

  • Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on visual environment and text instructions.

  • Multilingual Support: to serve global users, besides English and Chinese, Qwen2-VL now supports the understanding of texts in different languages inside images, including most European languages, Japanese, Korean, Arabic, Vietnamese, etc.

Model Architecture Updates:

Read the full model card (1,977 words)

Configuration

Architecture
Qwen2VLForConditionalGeneration
Context length (tokens)
32,768
Layers
28
Hidden size
3,584
Feed-forward size
18,944
Attention heads
28
Key/value heads
4
Vocabulary size
152,064
Sliding window (tokens)
32,768
RoPE base
1e+06
Stored precision
bfloat16
Model type
qwen2_vl

Identity and Version

Repository
Qwen/Qwen2-VL-7B-Instruct
Publisher
Qwen
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
8.3B parameters
Languages
en
Revision
eed13092ef92e448dd6875b2a00151bd3f7db0ac
First published
2024-08-28
Last updated
2025-02-06

Files and Weights

17 files, 16.6 GB in total. The weights are 5 files totalling 16.6 GB in safetensors.

Weights5 files · 16.6 GB
Configuration5 files · 59.3 KB
Tokenizer4 files · 11.5 MB
Documentation2 files · 29.1 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00005.safetensorsWeights3.9 GB eab4f4dc1abf
model-00002-of-00005.safetensorsWeights3.9 GB 0546b7cd070f
model-00003-of-00005.safetensorsWeights3.9 GB 11368ea1e9c0
model-00004-of-00005.safetensorsWeights3.9 GB 381e110d74db
model-00005-of-00005.safetensorsWeights1.1 GB e76ec4dc2e7c
chat_template.jsonConfiguration1.1 KB
config.jsonConfiguration1.2 KB
generation_config.jsonConfiguration244 B
model.safetensors.index.jsonConfiguration56.5 KB
preprocessor_config.jsonConfiguration347 B
LICENSEDocumentation11.3 KB
README.mdDocumentation17.7 KB
.gitattributesRepository1.5 KB
merges.txtTokenizer1.7 MB
tokenizer.jsonTokenizer7.0 MB
tokenizer_config.jsonTokenizer4.2 KB
vocab.jsonTokenizer2.8 MB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
16.6 GB
Download from Qwen

Released by Qwen through ModelScope. Read the license.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
likaixin/ScreenSpot-Pro Task android_studio_macosMetric android_studio_macosComparison conditions not established 0 ScreenSpot-Pro Leaderboard
Reported by a third party
Evaluated revision not stated 2026-03-17
likaixin/ScreenSpot-Pro Task eviews_windowsMetric eviews_windowsComparison conditions not established 12 ScreenSpot-Pro Leaderboard
Reported by a third party
Evaluated revision not stated 2026-03-17
likaixin/ScreenSpot-Pro Task fruitloops_windowsMetric fruitloops_windowsComparison conditions not established 1.8 ScreenSpot-Pro Leaderboard
Reported by a third party
Evaluated revision not stated 2026-03-17
likaixin/ScreenSpot-Pro Task linux_common_linuxMetric linux_common_linuxComparison conditions not established 2 ScreenSpot-Pro Leaderboard
Reported by a third party
Evaluated revision not stated 2026-03-17
likaixin/ScreenSpot-Pro Task matlab_macosMetric matlab_macosComparison conditions not established 2.2 ScreenSpot-Pro Leaderboard
Reported by a third party
Evaluated revision not stated 2026-03-17
likaixin/ScreenSpot-Pro Task overallMetric overallComparison conditions not established 1.6 ScreenSpot-Pro Leaderboard
Reported by a third party
Evaluated revision not stated 2026-03-17
likaixin/ScreenSpot-Pro Task powerpoint_windowsMetric powerpoint_windowsComparison conditions not established 2.4 ScreenSpot-Pro Leaderboard
Reported by a third party
Evaluated revision not stated 2026-03-17
likaixin/ScreenSpot-Pro Task unreal_engine_windowsMetric unreal_engine_windowsComparison conditions not established 2.9 ScreenSpot-Pro Leaderboard
Reported by a third party
Evaluated revision not stated 2026-03-17
likaixin/ScreenSpot-Pro Task vivado_windowsMetric vivado_windowsComparison conditions not established 1.2 ScreenSpot-Pro Leaderboard
Reported by a third party
Evaluated revision not stated 2026-03-17
likaixin/ScreenSpot-Pro Task vscode_macosMetric vscode_macosComparison conditions not established 5.5 ScreenSpot-Pro Leaderboard
Reported by a third party
Evaluated revision not stated 2026-03-17
likaixin/ScreenSpot-Pro Task word_macosMetric word_macosComparison conditions not established 6 ScreenSpot-Pro Leaderboard
Reported by a third party
Evaluated revision not stated 2026-03-17

Memory Requirements

PrecisionWeights in memory
As published16.6 GB
16-bit16.6 GB
8-bit8.3 GB
4-bit4.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Compare Qwen2-VL-7B-Instruct

Questions About Qwen2-VL-7B-Instruct

How much GPU memory does Qwen2-VL-7B-Instruct need?

About 19.9 GB at 16-bit and 5 GB at 4-bit: the weights (8.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen2-VL-7B-Instruct on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Qwen2-VL-7B-Instruct commercially?

Yes. Qwen2-VL-7B-Instruct is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Qwen2-VL-7B-Instruct's context length?

32,768 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen2-VL-7B-Instruct-AWQ

Qwen

We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation. SoTA understanding of images of various resolution & ratio: Qwen2-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc. Understanding videos of 20min+: Qwen2-VL can understand videos over 20 minutes for high-quality video-based question answering, dialog, content creation, etc. Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on…

Open weights apache-2.0 8.3B parameters 32,768 tokens transformers

Model · Image and text to text

Qwen2.5-VL-7B-Instruct

Qwen

pipelinetag: image-text-to-text - multimodal libraryname: transformers In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introduce the latest addition to the Qwen family: Qwen2.5-VL. Understanding long videos and capturing events: Qwen2.5-VL can comprehend videos of over 1 hour, and this time it has a new ability of cpaturing event by pinpointing the relevant video segments. Capable of visual localization in different formats: Qwen2.5-VL can accurately localize objects in…

Open weights apache-2.0 8.3B parameters 128,000 tokens transformers

Model · Image and text to text

Qwen2.5-VL-7B-Instruct-AWQ

Qwen

In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introduce the latest addition to the Qwen family: Qwen2.5-VL. Understanding long videos and capturing events: Qwen2.5-VL can comprehend videos of over 1 hour, and this time it has a new ability of cpaturing event by pinpointing the relevant video segments. Capable of visual localization in different formats: Qwen2.5-VL can accurately localize objects in an image by generating bounding boxes or points, and it can provide…

Open weights apache-2.0 8.3B parameters 128,000 tokens transformers

Model · Image and text to text

Qwen3-VL-8B-Instruct

Qwen

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities. Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning‑enhanced Thinking editions for flexible, on‑demand deployment. Text Understanding on par with pure LLMs: Seamless text–vision fusion for lossless, unified comprehension. 1. Interleaved-MRoPE: Full‑frequency allocation over time, width, and height…

Open weights apache-2.0 8.8B parameters 262,144 tokens transformers

Model · Image and text to text

MemGUI-8B-RL

Anonymous

Anonymous release for the ICLR 2027 submission MemGUI-RL: Reinforcement Learning for Proactive Context Management in Long-Horizon Mobile GUI Agents. Project page: https://memgui-rl-anonymous.github.io/ MemGUI-8B-RL is MemGUI-8B-SFT (Qwen3-VL-8B-Instruct supervised on MemGUI-3K) post-trained for 100 optimizer steps with FARPO (Folding-Aware Reward-decoupled Policy Optimization, span-to-step ratio rho = 9). The policy speaks the ConAct (Context-as-Action) interface of MemGUI-Agent: every response contains a folding directive for its own history, an optional memory operation and the next GUI action. The checkpoint is a standard Qwen3VLForConditionalGeneration model (weights in bf16, ~17.5 GB).…

Open weights apache-2.0 8.8B parameters 262,144 tokens

Model · Image and text to text

Qwen3-VL-8B-Instruct-FP8

Qwen

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities. Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning‑enhanced Thinking editions for flexible, on‑demand deployment. Text Understanding on par with pure LLMs: Seamless text–vision fusion for lossless, unified comprehension. 1. Interleaved-MRoPE: Full‑frequency allocation over time, width, and height…

Open weights apache-2.0 8.8B parameters 262,144 tokens transformers