This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Open-weight model · Image and text to text
DN-MOPD-Qwen3.5-2B-160updates
by XinLi XINLI1997/DN-MOPD-Qwen3.5-2B-160updates
DN-MOPD-Qwen3.5-2B-160updates is an open-weight model for image and text to text from XinLi, released under Apache License 2.0. It has 2.2B parameters and a 262,144-token context. At 16-bit it needs about 5.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
A Qwen3.5-2B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5).
Runs On
What it takes to serve DN-MOPD-Qwen3.5-2B-160updates (2.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 4.4 GB | 5.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 2.2 GB | 2.7 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 1.1 GB | 1.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.
DN-MOPD-Qwen3.5-2B-160updates on every accelerator the SAVRN Index prices, at every precision
Model Card
By XinLi, published under apache-2.0, revision e5763bc811b4.
A Qwen3.5-2B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5). Three same-size RL experts (math, code, instruction following) teach one student on its own responses; each prompt is scored by the expert of its domain, and DN-MOPD rescales each domain's token-level feedback by its measured spread, wd = clip(σall / σd, 0.25, 4), so that no domain dominates the shared update. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in…
Read XinLi's full model card
A Qwen3.5-2B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5). Three same-size RL experts (math, code, instruction following) teach one student on its own responses; each prompt is scored by the expert of its domain, and DN-MOPD rescales each domain's token-level feedback by its measured spread, w_d = clip(σ_all / σ_d, 0.25, 4), so that no domain dominates the shared update.
Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD
Model details
| Base model | Qwen/Qwen3.5-2B |
| Method | DN-MOPD: label-routed on-policy distillation with per-domain advantage scaling w_d = clip(σ_all / σ_d, 0.25, 4) |
| Teachers (same size) | math, code, IF |
| Training | 160 updates from the base model, student seed 42 |
| Precision | bfloat16 |
| Chat format | non-thinking (enable_thinking=False) |
| License | Apache-2.0 (same as the base model) |
Training recipe
- Prompts: 2,700 training prompts, 900 each for mathematics, code and instruction following; each prompt carries its domain label and is scored by that domain's expert (label routing).
- Advantage: per sampled token, teacher log-probability minus the actor-recomputed student log-probability, used in the clipped policy-gradient OPD loss (ratio clip 0.2/0.2); no KL or entropy term.
- DN-MOPD scaling: on every batch, σ_d is the population standard deviation of the teacher–rollout log-ratios over the valid response tokens of domain d, and σ_all pools all domains; each domain's advantages are multiplied by w_d = clip(σ_all / σ_d, 0.25, 4) (w_d = 1 if a statistic is degenerate). Signs are preserved.
- Batching: 64 prompts × 8 responses = 512 responses per update, one optimizer step per rollout batch.
- Lengths: prompt ≤ 2,048 tokens, response ≤ 8,192 tokens, temperature 1.0.
- Optimizer: Adam, learning rate 1e-6 (constant after 5 warm-up updates), betas (0.9, 0.98), weight decay 0.1, gradient clipping 1.0.
- Length of training: 160 updates (the 80-update run continued to 160), student seed 42.
The full recipe, with the launch scripts for every row of the paper's tables, is in
recipes/qwen3.5/ and docs/recipe.md.
Usage
This model was trained and evaluated with the non-thinking chat format. Pass enable_thinking=False to the chat
template. The evaluation settings in the paper were temperature 1.0 and top-p 1.0, with up to
16,384 new tokens (8,192 in the appendix).
vLLM (the paper used vLLM 0.18.0):
from vllm import LLM, SamplingParams
llm = LLM(model="XINLI1997/DN-MOPD-Qwen3.5-2B-160updates", max_model_len=32768)
params = SamplingParams(temperature=1.0, top_p=1.0, max_tokens=16384, seed=42)
messages = [{"role": "user", "content": "Find the sum of all positive divisors of 36. Put the final answer in \\boxed{}."}]
outputs = llm.chat(messages, params, chat_template_kwargs={"enable_thinking": False})
print(outputs[0].outputs[0].text)
Transformers (Qwen3.5 needs transformers>=5; the paper's training environment used 5.12.1):
import torch
from transformers import AutoModelForImageTextToText, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("XINLI1997/DN-MOPD-Qwen3.5-2B-160updates")
model = AutoModelForImageTextToText.from_pretrained("XINLI1997/DN-MOPD-Qwen3.5-2B-160updates", dtype=torch.bfloat16, device_map="auto")
messages = [{"role": "user", "content": "Write a Python function that returns the n-th Fibonacci number."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=4096, do_sample=True, temperature=1.0, top_p=1.0)
print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Evaluation
Paper Table 5 (six-task Total at 80 and 160 updates, 16K cap; changes use unrounded scores). This checkpoint is the 160-update endpoint of the bold row.
| Method (Qwen3.5-2B) | 80 | 160 | Δ Total |
|---|---|---|---|
| Single teacher (code) | 28.8 | 28.8 | -0.1 |
| Label-routed | 26.6 | 28.3 | +1.7 |
| DN-MOPD | 29.0 | 30.0 | +1.1 |
Scores (%) from the paper; training seed 42; 16,384-token evaluation cap; non-thinking chat template; temperature 1.0, top-p 1.0, generation seed 42. AIME25/AIME26: avg@64. LiveCodeBench v5/v6 (167/175 disjoint problems): avg@6. IFEval/IFBench: strict prompt accuracy, avg@16. Total: mean of the six task scores.
Files
- Weights in Hugging Face format (
Qwen3_5ForConditionalGeneration, bfloat16), exported from the FSDP training checkpoint. - The export omits the 15 multi-token-prediction tensors (
mtp.*) of the base model. All other tensors have the base model's names and shapes. MTP-based speculative decoding is therefore not available with this checkpoint. Ordinary decoding is unaffected: the paper's evaluations used exactly these files. config.json, the tokenizer files andchat_template.jinjaare the base model's, unchanged.- The vision encoder is carried over from the base model. Training and evaluation used text only.
LICENSEis the base model's Apache-2.0 license.
Limitations
- DN-MOPD is not the best integration recipe overall: in the paper, SeqKD-SFT and task-arithmetic merging (ParamMerge-TA) reach higher Totals under their own recipes. DN-MOPD's gains are over label-routed multi-teacher OPD and, by a smaller margin whose intervals sometimes include zero, over the strongest single-teacher student.
- Gains are largest in mathematics; code and IF gains are smaller and less consistent (at 9B, code does not improve at the 16K cap). The IF multiplier usually sits at the 0.25 lower bound, and the clipping bounds were not tuned per size.
- Trained on 2,700 prompts in three domains with a single seed (42) and one expert pool per size, within one model family; results for other teacher–student configurations may differ.
- Trained with responses of at most 8,192 tokens and evaluated only in non-thinking mode; thinking mode, multimodal inputs, other languages and safety behaviour were not evaluated beyond the base model.
Citation
@article{li2026dnmopd,
title = {Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation},
author = {Li, Xin and Jiang, Hao and Gao, Xin and Wang, Annan and Xie, Yuchen and Guo, Jinghao and Qu, Xingwei and Zhang, Yichi and Yuen, Chau},
journal = {arXiv preprint arXiv:2609.35347},
year = {2026},
url = {https://arxiv.org/abs/2609.35347}
}
This model is a fine-tuned derivative of Qwen/Qwen3.5-2B by the Qwen team, released under the Apache License 2.0.
Configuration
- Architecture
- Qwen3_5ForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 24
- Hidden size
- 2,048
- Feed-forward size
- 6,144
- Attention heads
- 8
- Key/value heads
- 2
- Head dimension
- 256
- Vocabulary size
- 248,320
- Model type
- qwen3_5
Identity and Version
- Repository
- XINLI1997/DN-MOPD-Qwen3.5-2B-160updates
- Publisher
- XinLi
- Task
- Image and text to text
- Modality
- Image and text
- Library
- transformers
- Parameters
- 2.2B parameters
- Languages
- en
- Revision
- e5763bc811b43ee2d1d2731e92aec07ff035cb9e
- First published
- 2026-10-01
- Last updated
- 2026-10-01
Files and Weights
13 files, 4.4 GB in total. The weights are 1 file totalling 4.4 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 4.4 GB | a61056536275 |
| config.json | Configuration | 2.9 KB | — |
| model.safetensors.index.json | Configuration | 45.2 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| video_preprocessor_config.json | Configuration | 385 B | — |
| LICENSE | Documentation | 11.5 KB | — |
| README.md | Documentation | 7.7 KB | — |
| chat_template.jinja | Other | 7.8 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| merges.txt | Tokenizer | 3.4 MB | — |
| tokenizer.json | Tokenizer | 12.8 MB | 5f9e4d4901a9 |
| tokenizer_config.json | Tokenizer | 16.7 KB | — |
| vocab.json | Tokenizer | 6.7 MB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 4.4 GB
Released by XinLi through its official repository on Hugging Face. Read the license.
Built From
- Derived from Qwen/Qwen3.5-2B
- Described by arXiv:2609.35347
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 4.4 GB |
| 16-bit | 4.4 GB |
| 8-bit | 2.2 GB |
| 4-bit | 1.1 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About DN-MOPD-Qwen3.5-2B-160updates
How much GPU memory does DN-MOPD-Qwen3.5-2B-160updates need?
About 5.3 GB at 16-bit and 1.3 GB at 4-bit: the weights (2.2B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run DN-MOPD-Qwen3.5-2B-160updates on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use DN-MOPD-Qwen3.5-2B-160updates commercially?
Yes. DN-MOPD-Qwen3.5-2B-160updates is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is DN-MOPD-Qwen3.5-2B-160updates's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.
Similar Models
The Label baseline of the DN-MOPD paper at Qwen3.5-2B continued to 160 updates (paper Table 5): multi-teacher on-policy distillation with label routing (each prompt is scored by the expert of its domain, every domain multiplier is 1). Released for comparison with DN-MOPD-Qwen3.5-2B; it is not the proposed method. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in recipes/qwen3.5/ and docs/recipe.md. This model was trained and evaluated with the non-thinking chat format. Pass enablethinking=False to the…
We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation. SoTA understanding of images of various resolution & ratio: Qwen2-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc. Understanding videos of 20min+: Qwen2-VL can understand videos over 20 minutes for high-quality video-based question answering, dialog, content creation, etc. Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on…
Model · Image and text to text
Qari-OCR-v0.3-VL-2B-Instruct
QARI-OCR v0.3 is a specialized vision-language model fine-tuned for Arabic Optical Character Recognition with a focus on structural document understanding. - Built on Qwen2-VL-2B-Instruct, this model excels at preserving document layouts, HTML tags, and formatting while transcribing Arabic text. - It is described in detail in the paper QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation. While QARI v0.2 achieves better raw text accuracy (CER: 0.061), QARI v0.3 excels in: - HTML/Markdown structure preservation - Document layout understanding - Handwritten text recognition (initial capabilities) - 5x faster training than v0.2 You can load this…
Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Scores of Qwen3.5 models are reported…
Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Rax 4.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. Rax 4.5 features the following enhancement: For more details, please refer to our blog post Rax 4.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not…
