The Label baseline of the DN-MOPD paper at Qwen3.5-4B continued to 160 updates (paper Table 5): multi-teacher on-policy distillation with label routing (each prompt is scored by the expert of its domain, every domain multiplier is 1). Released for comparison with DN-MOPD-Qwen3.5-4B; it is not the proposed method.
Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
(arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD
Model details
|
|
| Base model |
Qwen/Qwen3.5-4B |
| Method |
Label: label-routed multi-teacher on-policy distillation (MOPD), every w_d = 1 |
| Teachers (same size) |
math, code, IF |
| Training |
160 updates from the base model, student seed 42 |
| Precision |
bfloat16 |
| Chat format |
non-thinking (enable_thinking=False) |
| License |
Apache-2.0 (same as the base model) |
Training recipe
- Prompts: 2,700 training prompts, 900 each for mathematics, code and instruction following; each prompt carries its domain label and is scored by that domain's expert (label routing).
- Advantage: per sampled token, teacher log-probability minus the actor-recomputed student log-probability, used in the clipped policy-gradient OPD loss (ratio clip 0.2/0.2); no KL or entropy term.
- Label baseline: identical to DN-MOPD except that every domain multiplier is 1.
- Batching: 64 prompts × 8 responses = 512 responses per update, one optimizer step per rollout batch.
- Lengths: prompt ≤ 2,048 tokens, response ≤ 8,192 tokens, temperature 1.0.
- Optimizer: Adam, learning rate 1e-6 (constant after 5 warm-up updates), betas (0.9, 0.98), weight decay 0.1, gradient clipping 1.0.
- Length of training: 160 updates (the 80-update run continued to 160), student seed 42.
The full recipe, with the launch scripts for every row of the paper's tables, is in
recipes/qwen3.5/ and docs/recipe.md.
Usage
This model was trained and evaluated with the non-thinking chat format. Pass enable_thinking=False to the chat
template. Qwen3.5-4B's chat template enables thinking by default, so this argument is required. The evaluation settings in the paper were temperature 1.0 and top-p 1.0, with up to
16,384 new tokens (8,192 in the appendix).
vLLM (the paper used vLLM 0.18.0):
from vllm import LLM, SamplingParams
llm = LLM(model="XINLI1997/DN-MOPD-Qwen3.5-4B-baseline-label-160updates", max_model_len=32768)
params = SamplingParams(temperature=1.0, top_p=1.0, max_tokens=16384, seed=42)
messages = [{"role": "user", "content": "Find the sum of all positive divisors of 36. Put the final answer in \\boxed{}."}]
outputs = llm.chat(messages, params, chat_template_kwargs={"enable_thinking": False})
print(outputs[0].outputs[0].text)
Transformers (Qwen3.5 needs transformers>=5; the paper's training environment used 5.12.1):
import torch
from transformers import AutoModelForImageTextToText, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("XINLI1997/DN-MOPD-Qwen3.5-4B-baseline-label-160updates")
model = AutoModelForImageTextToText.from_pretrained("XINLI1997/DN-MOPD-Qwen3.5-4B-baseline-label-160updates", dtype=torch.bfloat16, device_map="auto")
messages = [{"role": "user", "content": "Write a Python function that returns the n-th Fibonacci number."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=4096, do_sample=True, temperature=1.0, top_p=1.0)
print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Evaluation
Paper Table 5 (six-task Total at 80 and 160 updates, 16K cap; changes use unrounded scores). This checkpoint is the 160-update endpoint of the bold row.
| Method (Qwen3.5-4B) |
80 |
160 |
Δ Total |
| Single teacher (code) |
50.9 |
51.2 |
+0.4 |
| Label-routed |
50.3 |
51.6 |
+1.3 |
| DN-MOPD |
52.5 |
53.1 |
+0.5 |
Scores (%) from the paper; training seed 42; 16,384-token evaluation cap; non-thinking chat template; temperature 1.0, top-p 1.0, generation seed 42. AIME25/AIME26: avg@64. LiveCodeBench v5/v6 (167/175 disjoint problems): avg@6. IFEval/IFBench: strict prompt accuracy, avg@16. Total: mean of the six task scores.
Files
- Weights in Hugging Face format (
Qwen3_5ForConditionalGeneration, bfloat16), exported from the FSDP training
checkpoint.
- The export omits the 15 multi-token-prediction tensors (
mtp.*) of the base model. All other tensors have the
base model's names and shapes. MTP-based speculative decoding is therefore not available with this checkpoint.
Ordinary decoding is unaffected: the paper's evaluations used exactly these files.
config.json, the tokenizer files and chat_template.jinja are the base model's, unchanged.
- The vision encoder is carried over from the base model. Training and evaluation used text only.
LICENSE is the base model's Apache-2.0 license.
Limitations
- This is the baseline, not the proposed method. In the paper it does not outperform the strongest single-teacher student at any size, and DN-MOPD improves on it at every size.
- Trained on 2,700 prompts in three domains with a single seed (42) and one expert pool per size, within one model family; results for other teacher–student configurations may differ.
- Trained with responses of at most 8,192 tokens and evaluated only in non-thinking mode; thinking mode, multimodal inputs, other languages and safety behaviour were not evaluated beyond the base model.
Citation
@article{li2026dnmopd,
title = {Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation},
author = {Li, Xin and Jiang, Hao and Gao, Xin and Wang, Annan and Xie, Yuchen and Guo, Jinghao and Qu, Xingwei and Zhang, Yichi and Yuen, Chau},
journal = {arXiv preprint arXiv:2609.35347},
year = {2026},
url = {https://arxiv.org/abs/2609.35347}
}
This model is a fine-tuned derivative of Qwen/Qwen3.5-4B by the Qwen team,
released under the Apache License 2.0.