SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

DN-MOPD-Qwen3.5-4B-160updates

by XinLi XINLI1997/DN-MOPD-Qwen3.5-4B-160updates

DN-MOPD-Qwen3.5-4B-160updates is an open-weight model for image and text to text from XinLi, released under Apache License 2.0. It has 4.5B parameters and a 262,144-token context. At 16-bit it needs about 10.9 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

A Qwen3.5-4B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5).

Parameters4.5B
Context262,144
Weights9.1 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve DN-MOPD-Qwen3.5-4B-160updates (4.5B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 9.1 GB 10.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 4.5 GB 5.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 2.3 GB 2.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

DN-MOPD-Qwen3.5-4B-160updates on every accelerator the SAVRN Index prices, at every precision

Model Card

By XinLi, published under apache-2.0, revision 123b7b4d6f25.

A Qwen3.5-4B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5). Three same-size RL experts (math, code, instruction following) teach one student on its own responses; each prompt is scored by the expert of its domain, and DN-MOPD rescales each domain's token-level feedback by its measured spread, wd = clip(σall / σd, 0.25, 4), so that no domain dominates the shared update. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in…

Read XinLi's full model card

A Qwen3.5-4B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5). Three same-size RL experts (math, code, instruction following) teach one student on its own responses; each prompt is scored by the expert of its domain, and DN-MOPD rescales each domain's token-level feedback by its measured spread, w_d = clip(σ_all / σ_d, 0.25, 4), so that no domain dominates the shared update.

Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD

Model details

Base model Qwen/Qwen3.5-4B
Method DN-MOPD: label-routed on-policy distillation with per-domain advantage scaling w_d = clip(σ_all / σ_d, 0.25, 4)
Teachers (same size) math, code, IF
Training 160 updates from the base model, student seed 42
Precision bfloat16
Chat format non-thinking (enable_thinking=False)
License Apache-2.0 (same as the base model)

Training recipe

  • Prompts: 2,700 training prompts, 900 each for mathematics, code and instruction following; each prompt carries its domain label and is scored by that domain's expert (label routing).
  • Advantage: per sampled token, teacher log-probability minus the actor-recomputed student log-probability, used in the clipped policy-gradient OPD loss (ratio clip 0.2/0.2); no KL or entropy term.
  • DN-MOPD scaling: on every batch, σ_d is the population standard deviation of the teacher–rollout log-ratios over the valid response tokens of domain d, and σ_all pools all domains; each domain's advantages are multiplied by w_d = clip(σ_all / σ_d, 0.25, 4) (w_d = 1 if a statistic is degenerate). Signs are preserved.
  • Batching: 64 prompts × 8 responses = 512 responses per update, one optimizer step per rollout batch.
  • Lengths: prompt ≤ 2,048 tokens, response ≤ 8,192 tokens, temperature 1.0.
  • Optimizer: Adam, learning rate 1e-6 (constant after 5 warm-up updates), betas (0.9, 0.98), weight decay 0.1, gradient clipping 1.0.
  • Length of training: 160 updates (the 80-update run continued to 160), student seed 42.

The full recipe, with the launch scripts for every row of the paper's tables, is in recipes/qwen3.5/ and docs/recipe.md.

Usage

This model was trained and evaluated with the non-thinking chat format. Pass enable_thinking=False to the chat template. Qwen3.5-4B's chat template enables thinking by default, so this argument is required. The evaluation settings in the paper were temperature 1.0 and top-p 1.0, with up to 16,384 new tokens (8,192 in the appendix).

vLLM (the paper used vLLM 0.18.0):

from vllm import LLM, SamplingParams

llm = LLM(model="XINLI1997/DN-MOPD-Qwen3.5-4B-160updates", max_model_len=32768)
params = SamplingParams(temperature=1.0, top_p=1.0, max_tokens=16384, seed=42)
messages = [{"role": "user", "content": "Find the sum of all positive divisors of 36. Put the final answer in \\boxed{}."}]
outputs = llm.chat(messages, params, chat_template_kwargs={"enable_thinking": False})
print(outputs[0].outputs[0].text)

Transformers (Qwen3.5 needs transformers>=5; the paper's training environment used 5.12.1):

import torch
from transformers import AutoModelForImageTextToText, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("XINLI1997/DN-MOPD-Qwen3.5-4B-160updates")
model = AutoModelForImageTextToText.from_pretrained("XINLI1997/DN-MOPD-Qwen3.5-4B-160updates", dtype=torch.bfloat16, device_map="auto")
messages = [{"role": "user", "content": "Write a Python function that returns the n-th Fibonacci number."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=4096, do_sample=True, temperature=1.0, top_p=1.0)
print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Evaluation

Paper Table 5 (six-task Total at 80 and 160 updates, 16K cap; changes use unrounded scores). This checkpoint is the 160-update endpoint of the bold row.

Method (Qwen3.5-4B) 80 160 Δ Total
Single teacher (code) 50.9 51.2 +0.4
Label-routed 50.3 51.6 +1.3
DN-MOPD 52.5 53.1 +0.5

Scores (%) from the paper; training seed 42; 16,384-token evaluation cap; non-thinking chat template; temperature 1.0, top-p 1.0, generation seed 42. AIME25/AIME26: avg@64. LiveCodeBench v5/v6 (167/175 disjoint problems): avg@6. IFEval/IFBench: strict prompt accuracy, avg@16. Total: mean of the six task scores.

Files

  • Weights in Hugging Face format (Qwen3_5ForConditionalGeneration, bfloat16), exported from the FSDP training checkpoint.
  • The export omits the 15 multi-token-prediction tensors (mtp.*) of the base model. All other tensors have the base model's names and shapes. MTP-based speculative decoding is therefore not available with this checkpoint. Ordinary decoding is unaffected: the paper's evaluations used exactly these files.
  • config.json, the tokenizer files and chat_template.jinja are the base model's, unchanged.
  • The vision encoder is carried over from the base model. Training and evaluation used text only.
  • LICENSE is the base model's Apache-2.0 license.

Limitations

  • DN-MOPD is not the best integration recipe overall: in the paper, SeqKD-SFT and task-arithmetic merging (ParamMerge-TA) reach higher Totals under their own recipes. DN-MOPD's gains are over label-routed multi-teacher OPD and, by a smaller margin whose intervals sometimes include zero, over the strongest single-teacher student.
  • Gains are largest in mathematics; code and IF gains are smaller and less consistent (at 9B, code does not improve at the 16K cap). The IF multiplier usually sits at the 0.25 lower bound, and the clipping bounds were not tuned per size.
  • Trained on 2,700 prompts in three domains with a single seed (42) and one expert pool per size, within one model family; results for other teacher–student configurations may differ.
  • Trained with responses of at most 8,192 tokens and evaluated only in non-thinking mode; thinking mode, multimodal inputs, other languages and safety behaviour were not evaluated beyond the base model.

Citation

@article{li2026dnmopd,
  title   = {Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation},
  author  = {Li, Xin and Jiang, Hao and Gao, Xin and Wang, Annan and Xie, Yuchen and Guo, Jinghao and Qu, Xingwei and Zhang, Yichi and Yuen, Chau},
  journal = {arXiv preprint arXiv:2609.35347},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.35347}
}

This model is a fine-tuned derivative of Qwen/Qwen3.5-4B by the Qwen team, released under the Apache License 2.0.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
32
Hidden size
2,560
Feed-forward size
9,216
Attention heads
16
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
XINLI1997/DN-MOPD-Qwen3.5-4B-160updates
Publisher
XinLi
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
4.5B parameters
Languages
en
Revision
123b7b4d6f259c0450557c8615903c5170d78464
First published
2026-10-01
Last updated
2026-10-01

Files and Weights

14 files, 9.1 GB in total. The weights are 2 files totalling 9.1 GB in safetensors.

Weights2 files · 9.1 GB
Configuration4 files · 70.1 KB
Tokenizer4 files · 22.9 MB
Documentation2 files · 19.4 KB
Other1 file · 7.8 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00002.safetensorsWeights5.0 GB 42d046555ee6
model-00002-of-00002.safetensorsWeights4.1 GB 773c45668227
config.jsonConfiguration3.2 KB —
model.safetensors.index.jsonConfiguration66.2 KB —
preprocessor_config.jsonConfiguration390 B —
video_preprocessor_config.jsonConfiguration385 B —
LICENSEDocumentation11.5 KB —
README.mdDocumentation7.8 KB —
chat_template.jinjaOther7.8 KB —
.gitattributesRepository1.6 KB —
merges.txtTokenizer3.4 MB —
tokenizer.jsonTokenizer12.8 MB 5f9e4d4901a9
tokenizer_config.jsonTokenizer16.7 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
9.1 GB
Download from XinLi

Released by XinLi through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published9.1 GB
16-bit9.1 GB
8-bit4.5 GB
4-bit2.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About DN-MOPD-Qwen3.5-4B-160updates

How much GPU memory does DN-MOPD-Qwen3.5-4B-160updates need?

About 10.9 GB at 16-bit and 2.7 GB at 4-bit: the weights (4.5B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run DN-MOPD-Qwen3.5-4B-160updates on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use DN-MOPD-Qwen3.5-4B-160updates commercially?

Yes. DN-MOPD-Qwen3.5-4B-160updates is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is DN-MOPD-Qwen3.5-4B-160updates's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Vinci-Piccolo-1.0

SimpleDirect

Vinci Piccolo is a small, open-weight chat model fine-tuned for character and honesty — the first model in the Vinci family from SimpleDirect. The character you'd want in an AI, open and small enough to run yourself. Try it: chat app — free · ollama run hf.co/simpledirect/Vinci-Piccolo-1.0-GGUF Weight-file size is not a runtime-memory requirement — model loading, KV cache, context length and batching all need memory beyond the weights. See Hardware requirements below for the figures we do give. Most fine-tuning optimizes for capability. Vinci Piccolo is fine-tuned for something else: a consistent character and an honest disposition. It is trained against a written, public Constitution that…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

Omni-Edu-4B

Hao Liang

This model is a fine-tuned version of Qwen/Qwen3.5-4B-Base on the Omni-Edu-70K dataset. The following hyperparameters were used during training: - learningrate: 5e-06 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 8 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 64 - totalevalbatchsize: 64 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 3.0 - Transformers 5.2.0 - Pytorch 2.10.0 - Datasets 4.0.0 - Tokenizers 0.22.2

Open weights other 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

DN-MOPD-Qwen3.5-4B-baseline-label-160updates

XinLi

The Label baseline of the DN-MOPD paper at Qwen3.5-4B continued to 160 updates (paper Table 5): multi-teacher on-policy distillation with label routing (each prompt is scored by the expert of its domain, every domain multiplier is 1). Released for comparison with DN-MOPD-Qwen3.5-4B; it is not the proposed method. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in recipes/qwen3.5/ and docs/recipe.md. This model was trained and evaluated with the non-thinking chat format. Pass enablethinking=False to the…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3-VL-4B-Instruct

Qwen

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities. Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning‑enhanced Thinking editions for flexible, on‑demand deployment. Text Understanding on par with pure LLMs: Seamless text–vision fusion for lossless, unified comprehension. 1. Interleaved-MRoPE: Full‑frequency allocation over time, width, and height…

Open weights apache-2.0 4.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.5-4B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 4.7B parameters 262,144 tokens transformers

Model · Image and text to text

Lodestar-4B

StartLux

Lodestar-4B is a 4-billion-parameter decision model. You give it a state (plain text, JSON or a long document) and one or more typed questions; it returns a probability for every listed option. Each answer is read from a single forward pass. The model never generates free text, so there is nothing to parse and no output length to budget for. (Probabilities rounded to three decimals.) The same call from Python, run inside the downloaded folder: A score question takes its levels as a list, lowest first, and returns {"type": "score", "score":, "probabilities": {"0": p0, "1": p1,...}}. Every question is rendered into one chat prompt (thinking disabled): The probability of each option is the…

Open weights apache-2.0 4.7B parameters 262,144 tokens transformers