SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Qwen3.5-2B-Vi-SFT-VLM-v2.1

by Pham Phuc Phuc-HugigFace/Qwen3.5-2B-Vi-SFT-VLM-v2.1

Qwen3.5-2B-Vi-SFT-VLM-v2.1 is an open-weight model for image and text to text from Pham Phuc. It has 2.2B parameters and a 262,144-token context. At 16-bit it needs about 5.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model.

Parameters2.2B
Context262,144
Weights4.4 GB
License—
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Qwen3.5-2B-Vi-SFT-VLM-v2.1 (2.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 4.4 GB 5.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 2.2 GB 2.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 1.1 GB 1.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

Qwen3.5-2B-Vi-SFT-VLM-v2.1 on every accelerator the SAVRN Index prices, at every precision

Model Card

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Excerpt from the card by Pham Phuc.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
24
Hidden size
2,048
Feed-forward size
6,144
Attention heads
8
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
Phuc-HugigFace/Qwen3.5-2B-Vi-SFT-VLM-v2.1
Publisher
Pham Phuc
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
2.2B parameters
Languages
Not stated by the source
Revision
85b57ff57f5fa0082a81dbc180465104cbadf22f
First published
2026-09-24
Last updated
2026-09-24

Files and Weights

9 files, 4.4 GB in total. The weights are 1 file totalling 4.4 GB in safetensors.

Weights1 file · 4.4 GB
Configuration3 files · 4.0 KB
Tokenizer2 files · 20.0 MB
Documentation1 file · 5.2 KB
Other1 file · 7.8 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights4.4 GB 155a44fd423a
config.jsonConfiguration2.7 KB —
generation_config.jsonConfiguration116 B —
processor_config.jsonConfiguration1.2 KB —
README.mdDocumentation5.2 KB —
chat_template.jinjaOther7.8 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.1 KB —

License and Download

License
Not stated by the source
Access
Open weights, no gate
Download size
4.4 GB
Download from Pham Phuc

Released by Pham Phuc through its official repository on Hugging Face.

Built From

Memory Requirements

PrecisionWeights in memory
As published4.4 GB
16-bit4.4 GB
8-bit2.2 GB
4-bit1.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3.5-2B-Vi-SFT-VLM-v2.1

How much GPU memory does Qwen3.5-2B-Vi-SFT-VLM-v2.1 need?

About 5.3 GB at 16-bit and 1.3 GB at 4-bit: the weights (2.2B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Qwen3.5-2B-Vi-SFT-VLM-v2.1 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What is Qwen3.5-2B-Vi-SFT-VLM-v2.1's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

DN-MOPD-Qwen3.5-2B-baseline-label-160updates

XinLi

The Label baseline of the DN-MOPD paper at Qwen3.5-2B continued to 160 updates (paper Table 5): multi-teacher on-policy distillation with label routing (each prompt is scored by the expert of its domain, every domain multiplier is 1). Released for comparison with DN-MOPD-Qwen3.5-2B; it is not the proposed method. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in recipes/qwen3.5/ and docs/recipe.md. This model was trained and evaluated with the non-thinking chat format. Pass enablethinking=False to the…

Open weights apache-2.0 2.2B parameters 262,144 tokens transformers

Model · Image and text to text

DN-MOPD-Qwen3.5-2B-160updates

XinLi

A Qwen3.5-2B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5). Three same-size RL experts (math, code, instruction following) teach one student on its own responses; each prompt is scored by the expert of its domain, and DN-MOPD rescales each domain's token-level feedback by its measured spread, wd = clip(σall / σd, 0.25, 4), so that no domain dominates the shared update. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in…

Open weights apache-2.0 2.2B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen2-VL-2B-Instruct

Qwen

We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation. SoTA understanding of images of various resolution & ratio: Qwen2-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc. Understanding videos of 20min+: Qwen2-VL can understand videos over 20 minutes for high-quality video-based question answering, dialog, content creation, etc. Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on…

Open weights apache-2.0 2.2B parameters 32,768 tokens transformers

QARI-OCR v0.3 is a specialized vision-language model fine-tuned for Arabic Optical Character Recognition with a focus on structural document understanding. - Built on Qwen2-VL-2B-Instruct, this model excels at preserving document layouts, HTML tags, and formatting while transcribing Arabic text. - It is described in detail in the paper QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation. While QARI v0.2 achieves better raw text accuracy (CER: 0.061), QARI v0.3 excels in: - HTML/Markdown structure preservation - Document layout understanding - Handwritten text recognition (initial capabilities) - 5x faster training than v0.2 You can load this…

Open weights apache-2.0 2.2B parameters 32,768 tokens transformers

Model · Image and text to text

Qwen3.5-2B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Scores of Qwen3.5 models are reported…

Open weights apache-2.0 2.3B parameters 262,144 tokens transformers

Model · Image and text to text

Rax-4.5

RaxCore

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Rax 4.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. Rax 4.5 features the following enhancement: For more details, please refer to our blog post Rax 4.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not…

Open weights apache-2.0 2.3B parameters 262,144 tokens transformers