Anonymous release for the ICLR 2027 submission MemGUI-RL: Reinforcement Learning for Proactive Context Management in Long-Horizon Mobile GUI Agents. Project page: https://memgui-rl-anonymous.github.io/ MemGUI-8B-RL is MemGUI-8B-SFT (Qwen3-VL-8B-Instruct supervised on MemGUI-3K) post-trained for 100 optimizer steps with FARPO (Folding-Aware Reward-decoupled Policy Optimization, span-to-step ratio rho = 9). The policy speaks the ConAct (Context-as-Action) interface of MemGUI-Agent: every response contains a folding directive for its own history, an optional memory operation and the next GUI action. The checkpoint is a standard Qwen3VLForConditionalGeneration model (weights in bf16, ~17.5 GB).…
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
Runs On
What it takes to serve Qwen3-VL-8B-Instruct (8.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 17.5 GB | 21.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 8.8 GB | 10.5 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 4.4 GB | 5.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.
Model Card
By Qwen, published under apache-2.0, revision 0c351dd01ed8.
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities.
Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning‑enhanced Thinking editions for flexible, on‑demand deployment.
Key Enhancements:
Configuration
- Architecture
- Qwen3VLForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 36
- Hidden size
- 4,096
- Feed-forward size
- 12,288
- Attention heads
- 32
- Key/value heads
- 8
- Head dimension
- 128
- Vocabulary size
- 151,936
- RoPE base
- 5,000,000
- Model type
- qwen3_vl
Identity and Version
- Repository
- Qwen/Qwen3-VL-8B-Instruct
- Publisher
- Qwen
- Task
- Image and text to text
- Modality
- Image and text
- Library
- transformers
- Parameters
- 8.8B parameters
- Languages
- Not stated by the source
- Revision
- 0c351dd01ed87e9c1b53cbc748cba10e6187ff3b
- First published
- 2025-10-11
- Last updated
- 2025-10-15
Files and Weights
16 files, 17.5 GB in total. The weights are 4 files totalling 17.5 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00004.safetensors | Weights | 4.9 GB | d5d0aef0eb17 |
| model-00002-of-00004.safetensors | Weights | 4.9 GB | 8be88fb5501e |
| model-00003-of-00004.safetensors | Weights | 5.0 GB | 83de00eafe6e |
| model-00004-of-00004.safetensors | Weights | 2.7 GB | 0a88b98e9f96 |
| chat_template.json | Configuration | 5.5 KB | — |
| config.json | Configuration | 1.5 KB | — |
| generation_config.json | Configuration | 269 B | — |
| model.safetensors.index.json | Configuration | 67.8 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| video_preprocessor_config.json | Configuration | 385 B | — |
| README.md | Documentation | 7.1 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| merges.txt | Tokenizer | 1.7 MB | — |
| tokenizer.json | Tokenizer | 7.0 MB | — |
| tokenizer_config.json | Tokenizer | 10.9 KB | — |
| vocab.json | Tokenizer | 2.8 MB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 17.5 GB
Released by Qwen through ModelScope. Read the license.
Built From
- Described by arXiv:2308.12966
- Described by arXiv:2409.12191
- Described by arXiv:2502.13923
- Described by arXiv:2505.09388
Evaluations
Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| Delores-Lin/MDPBench | Task arMetric arComparison conditions not established | 63.1 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task deMetric deComparison conditions not established | 73.7 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task digitalMetric digitalComparison conditions not established | 78.4 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task enMetric enComparison conditions not established | 71.4 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task esMetric esComparison conditions not established | 69.3 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task frMetric frComparison conditions not established | 66.2 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task hiMetric hiComparison conditions not established | 58.4 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task idMetric idComparison conditions not established | 68.5 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task itMetric itComparison conditions not established | 79.1 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task jpMetric jpComparison conditions not established | 59.9 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task koMetric koComparison conditions not established | 61.9 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task latinMetric latinComparison conditions not established | 73.6 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task nlMetric nlComparison conditions not established | 78.3 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task non_latinMetric non_latinComparison conditions not established | 62.5 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task overallMetric overallComparison conditions not established | 68.3 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task photographedMetric photographedComparison conditions not established | 65 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task ptMetric ptComparison conditions not established | 82.2 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task ruMetric ruComparison conditions not established | 57.9 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task thMetric thComparison conditions not established | 62 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task viMetric viComparison conditions not established | 73.4 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task zhMetric zhComparison conditions not established | 62.6 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| Delores-Lin/MDPBench | Task zh_tMetric zh_tComparison conditions not established | 73.8 | MDPBench leaderboard Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| llamaindex/ParseBench | Task chartMetric chartSetup Pipeline name: qwen3vl_layoutComparison conditions not established | 1.4 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| llamaindex/ParseBench | Task layoutMetric layoutSetup Pipeline name: qwen3vl_layoutComparison conditions not established | 55.1 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| llamaindex/ParseBench | Task meanMetric meanSetup Pipeline name: qwen3vl_layoutComparison conditions not established | 46.8 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| llamaindex/ParseBench | Task tableMetric tableSetup Pipeline name: qwen3vl_layoutComparison conditions not established | 21.4 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| llamaindex/ParseBench | Task text_contentMetric text_contentSetup Pipeline name: qwen3vl_layoutComparison conditions not established | 89.5 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| llamaindex/ParseBench | Task text_formattingMetric text_formattingSetup Pipeline name: qwen3vl_layoutComparison conditions not established | 66.8 | ParseBench Reported by a third party |
Evaluated revision not stated | 2026-04-14 |
| tiiuae/PBench | Task averageMetric averageSetup Detection-only model combined with SAM2 to convert boxes to segmentation masks.Comparison conditions not established | 49 | Community Evals Reported by a third party |
Evaluated revision not stated | 2026-05-11 |
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 17.5 GB |
| 16-bit | 17.5 GB |
| 8-bit | 8.8 GB |
| 4-bit | 4.4 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Built on This Model
- Quantized fromQwen3-VL-8B-Instruct-FP8
- Derived fromQwen3-VL-8B-Instruct-FP8
- Derived fromQwen3-VL-Embedding-8B
- Quantized fromQwen3-VL-8B-Instruct-NVFP4
- Derived fromQwen3-VL-8B-Instruct-NVFP4
Compare Qwen3-VL-8B-Instruct
Questions About Qwen3-VL-8B-Instruct
How much GPU memory does Qwen3-VL-8B-Instruct need?
About 21 GB at 16-bit and 5.3 GB at 4-bit: the weights (8.8B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Qwen3-VL-8B-Instruct on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use Qwen3-VL-8B-Instruct commercially?
Yes. Qwen3-VL-8B-Instruct is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is Qwen3-VL-8B-Instruct's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.
Similar Models
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities. Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning‑enhanced Thinking editions for flexible, on‑demand deployment. Text Understanding on par with pure LLMs: Seamless text–vision fusion for lossless, unified comprehension. 1. Interleaved-MRoPE: Full‑frequency allocation over time, width, and height…
pipelinetag: image-text-to-text - multimodal libraryname: transformers In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introduce the latest addition to the Qwen family: Qwen2.5-VL. Understanding long videos and capturing events: Qwen2.5-VL can comprehend videos of over 1 hour, and this time it has a new ability of cpaturing event by pinpointing the relevant video segments. Capable of visual localization in different formats: Qwen2.5-VL can accurately localize objects in…
In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introduce the latest addition to the Qwen family: Qwen2.5-VL. Understanding long videos and capturing events: Qwen2.5-VL can comprehend videos of over 1 hour, and this time it has a new ability of cpaturing event by pinpointing the relevant video segments. Capable of visual localization in different formats: Qwen2.5-VL can accurately localize objects in an image by generating bounding boxes or points, and it can provide…
We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation. SoTA understanding of images of various resolution & ratio: Qwen2-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc. Understanding videos of 20min+: Qwen2-VL can understand videos over 20 minutes for high-quality video-based question answering, dialog, content creation, etc. Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on…
We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation. SoTA understanding of images of various resolution & ratio: Qwen2-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc. Understanding videos of 20min+: Qwen2-VL can understand videos over 20 minutes for high-quality video-based question answering, dialog, content creation, etc. Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on…