SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

DN-MOPD-Qwen3.5-2B-160updates

by XinLi XINLI1997/DN-MOPD-Qwen3.5-2B-160updates

DN-MOPD-Qwen3.5-2B-160updates is an open-weight model for image and text to text from XinLi, released under Apache License 2.0. It has 2.2B parameters and a 262,144-token context. At 16-bit it needs about 5.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

A Qwen3.5-2B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5).

Parameters2.2B
Context262,144
Weights4.4 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve DN-MOPD-Qwen3.5-2B-160updates (2.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 4.4 GB 5.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 2.2 GB 2.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 1.1 GB 1.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

DN-MOPD-Qwen3.5-2B-160updates on every accelerator the SAVRN Index prices, at every precision

Model Card

By XinLi, published under apache-2.0, revision e5763bc811b4.

A Qwen3.5-2B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5). Three same-size RL experts (math, code, instruction following) teach one student on its own responses; each prompt is scored by the expert of its domain, and DN-MOPD rescales each domain's token-level feedback by its measured spread, wd = clip(σall / σd, 0.25, 4), so that no domain dominates the shared update. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in…

Read XinLi's full model card

A Qwen3.5-2B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5). Three same-size RL experts (math, code, instruction following) teach one student on its own responses; each prompt is scored by the expert of its domain, and DN-MOPD rescales each domain's token-level feedback by its measured spread, w_d = clip(σ_all / σ_d, 0.25, 4), so that no domain dominates the shared update.

Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD

Model details

Base model Qwen/Qwen3.5-2B
Method DN-MOPD: label-routed on-policy distillation with per-domain advantage scaling w_d = clip(σ_all / σ_d, 0.25, 4)
Teachers (same size) math, code, IF
Training 160 updates from the base model, student seed 42
Precision bfloat16
Chat format non-thinking (enable_thinking=False)
License Apache-2.0 (same as the base model)

Training recipe

  • Prompts: 2,700 training prompts, 900 each for mathematics, code and instruction following; each prompt carries its domain label and is scored by that domain's expert (label routing).
  • Advantage: per sampled token, teacher log-probability minus the actor-recomputed student log-probability, used in the clipped policy-gradient OPD loss (ratio clip 0.2/0.2); no KL or entropy term.
  • DN-MOPD scaling: on every batch, σ_d is the population standard deviation of the teacher–rollout log-ratios over the valid response tokens of domain d, and σ_all pools all domains; each domain's advantages are multiplied by w_d = clip(σ_all / σ_d, 0.25, 4) (w_d = 1 if a statistic is degenerate). Signs are preserved.
  • Batching: 64 prompts × 8 responses = 512 responses per update, one optimizer step per rollout batch.
  • Lengths: prompt ≤ 2,048 tokens, response ≤ 8,192 tokens, temperature 1.0.
  • Optimizer: Adam, learning rate 1e-6 (constant after 5 warm-up updates), betas (0.9, 0.98), weight decay 0.1, gradient clipping 1.0.
  • Length of training: 160 updates (the 80-update run continued to 160), student seed 42.

The full recipe, with the launch scripts for every row of the paper's tables, is in recipes/qwen3.5/ and docs/recipe.md.

Usage

This model was trained and evaluated with the non-thinking chat format. Pass enable_thinking=False to the chat template. The evaluation settings in the paper were temperature 1.0 and top-p 1.0, with up to 16,384 new tokens (8,192 in the appendix).

vLLM (the paper used vLLM 0.18.0):

from vllm import LLM, SamplingParams

llm = LLM(model="XINLI1997/DN-MOPD-Qwen3.5-2B-160updates", max_model_len=32768)
params = SamplingParams(temperature=1.0, top_p=1.0, max_tokens=16384, seed=42)
messages = [{"role": "user", "content": "Find the sum of all positive divisors of 36. Put the final answer in \\boxed{}."}]
outputs = llm.chat(messages, params, chat_template_kwargs={"enable_thinking": False})
print(outputs[0].outputs[0].text)

Transformers (Qwen3.5 needs transformers>=5; the paper's training environment used 5.12.1):

import torch
from transformers import AutoModelForImageTextToText, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("XINLI1997/DN-MOPD-Qwen3.5-2B-160updates")
model = AutoModelForImageTextToText.from_pretrained("XINLI1997/DN-MOPD-Qwen3.5-2B-160updates", dtype=torch.bfloat16, device_map="auto")
messages = [{"role": "user", "content": "Write a Python function that returns the n-th Fibonacci number."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=4096, do_sample=True, temperature=1.0, top_p=1.0)
print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Evaluation

Paper Table 5 (six-task Total at 80 and 160 updates, 16K cap; changes use unrounded scores). This checkpoint is the 160-update endpoint of the bold row.

Method (Qwen3.5-2B) 80 160 Δ Total
Single teacher (code) 28.8 28.8 -0.1
Label-routed 26.6 28.3 +1.7
DN-MOPD 29.0 30.0 +1.1

Scores (%) from the paper; training seed 42; 16,384-token evaluation cap; non-thinking chat template; temperature 1.0, top-p 1.0, generation seed 42. AIME25/AIME26: avg@64. LiveCodeBench v5/v6 (167/175 disjoint problems): avg@6. IFEval/IFBench: strict prompt accuracy, avg@16. Total: mean of the six task scores.

Files

  • Weights in Hugging Face format (Qwen3_5ForConditionalGeneration, bfloat16), exported from the FSDP training checkpoint.
  • The export omits the 15 multi-token-prediction tensors (mtp.*) of the base model. All other tensors have the base model's names and shapes. MTP-based speculative decoding is therefore not available with this checkpoint. Ordinary decoding is unaffected: the paper's evaluations used exactly these files.
  • config.json, the tokenizer files and chat_template.jinja are the base model's, unchanged.
  • The vision encoder is carried over from the base model. Training and evaluation used text only.
  • LICENSE is the base model's Apache-2.0 license.

Limitations

  • DN-MOPD is not the best integration recipe overall: in the paper, SeqKD-SFT and task-arithmetic merging (ParamMerge-TA) reach higher Totals under their own recipes. DN-MOPD's gains are over label-routed multi-teacher OPD and, by a smaller margin whose intervals sometimes include zero, over the strongest single-teacher student.
  • Gains are largest in mathematics; code and IF gains are smaller and less consistent (at 9B, code does not improve at the 16K cap). The IF multiplier usually sits at the 0.25 lower bound, and the clipping bounds were not tuned per size.
  • Trained on 2,700 prompts in three domains with a single seed (42) and one expert pool per size, within one model family; results for other teacher–student configurations may differ.
  • Trained with responses of at most 8,192 tokens and evaluated only in non-thinking mode; thinking mode, multimodal inputs, other languages and safety behaviour were not evaluated beyond the base model.

Citation

@article{li2026dnmopd,
  title   = {Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation},
  author  = {Li, Xin and Jiang, Hao and Gao, Xin and Wang, Annan and Xie, Yuchen and Guo, Jinghao and Qu, Xingwei and Zhang, Yichi and Yuen, Chau},
  journal = {arXiv preprint arXiv:2609.35347},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.35347}
}

This model is a fine-tuned derivative of Qwen/Qwen3.5-2B by the Qwen team, released under the Apache License 2.0.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
24
Hidden size
2,048
Feed-forward size
6,144
Attention heads
8
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
XINLI1997/DN-MOPD-Qwen3.5-2B-160updates
Publisher
XinLi
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
2.2B parameters
Languages
en
Revision
e5763bc811b43ee2d1d2731e92aec07ff035cb9e
First published
2026-10-01
Last updated
2026-10-01

Files and Weights

13 files, 4.4 GB in total. The weights are 1 file totalling 4.4 GB in safetensors.

Weights1 file · 4.4 GB
Configuration4 files · 48.9 KB
Tokenizer4 files · 22.9 MB
Documentation2 files · 19.3 KB
Other1 file · 7.8 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights4.4 GB a61056536275
config.jsonConfiguration2.9 KB —
model.safetensors.index.jsonConfiguration45.2 KB —
preprocessor_config.jsonConfiguration390 B —
video_preprocessor_config.jsonConfiguration385 B —
LICENSEDocumentation11.5 KB —
README.mdDocumentation7.7 KB —
chat_template.jinjaOther7.8 KB —
.gitattributesRepository1.6 KB —
merges.txtTokenizer3.4 MB —
tokenizer.jsonTokenizer12.8 MB 5f9e4d4901a9
tokenizer_config.jsonTokenizer16.7 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
4.4 GB
Download from XinLi

Released by XinLi through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published4.4 GB
16-bit4.4 GB
8-bit2.2 GB
4-bit1.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About DN-MOPD-Qwen3.5-2B-160updates

How much GPU memory does DN-MOPD-Qwen3.5-2B-160updates need?

About 5.3 GB at 16-bit and 1.3 GB at 4-bit: the weights (2.2B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run DN-MOPD-Qwen3.5-2B-160updates on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use DN-MOPD-Qwen3.5-2B-160updates commercially?

Yes. DN-MOPD-Qwen3.5-2B-160updates is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is DN-MOPD-Qwen3.5-2B-160updates's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.5-2B-Vi-SFT-VLM-v2.1

Pham Phuc

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Open weights 2.2B parameters 262,144 tokens transformers

Model · Image and text to text

DN-MOPD-Qwen3.5-2B-baseline-label-160updates

XinLi

The Label baseline of the DN-MOPD paper at Qwen3.5-2B continued to 160 updates (paper Table 5): multi-teacher on-policy distillation with label routing (each prompt is scored by the expert of its domain, every domain multiplier is 1). Released for comparison with DN-MOPD-Qwen3.5-2B; it is not the proposed method. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in recipes/qwen3.5/ and docs/recipe.md. This model was trained and evaluated with the non-thinking chat format. Pass enablethinking=False to the…

Open weights apache-2.0 2.2B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen2-VL-2B-Instruct

Qwen

We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation. SoTA understanding of images of various resolution & ratio: Qwen2-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc. Understanding videos of 20min+: Qwen2-VL can understand videos over 20 minutes for high-quality video-based question answering, dialog, content creation, etc. Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on…

Open weights apache-2.0 2.2B parameters 32,768 tokens transformers

QARI-OCR v0.3 is a specialized vision-language model fine-tuned for Arabic Optical Character Recognition with a focus on structural document understanding. - Built on Qwen2-VL-2B-Instruct, this model excels at preserving document layouts, HTML tags, and formatting while transcribing Arabic text. - It is described in detail in the paper QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation. While QARI v0.2 achieves better raw text accuracy (CER: 0.061), QARI v0.3 excels in: - HTML/Markdown structure preservation - Document layout understanding - Handwritten text recognition (initial capabilities) - 5x faster training than v0.2 You can load this…

Open weights apache-2.0 2.2B parameters 32,768 tokens transformers

Model · Image and text to text

Qwen3.5-2B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Scores of Qwen3.5 models are reported…

Open weights apache-2.0 2.3B parameters 262,144 tokens transformers

Model · Image and text to text

Rax-4.5

RaxCore

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Rax 4.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. Rax 4.5 features the following enhancement: For more details, please refer to our blog post Rax 4.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not…

Open weights apache-2.0 2.3B parameters 262,144 tokens transformers