SAVRN
Search Contact SAVRN

Research paper · 2026-09-28

Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

Xin Li, Hao Jiang, Xin Gao, Annan Wang, Yuchen Xie, Jinghao Guo, Xingwei Qu, Yichi Zhang, Chau Yuen

5 open models in the SAVRN Model Hub cite Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (2026). The most downloaded is DN-MOPD-Qwen3.5-2B-baseline-label-160updates by XinLi (image and text to text, 2.2B parameters).

Published2026-09-28
Authors9
Citing Models5
arXiv2609.35347

Abstract

Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.

Full paper on arXiv · Code

Details

arXiv identifier
2609.35347
Published
2026-09-28
Authors
Xin Li, Hao Jiang, Xin Gao, Annan Wang, Yuchen Xie, Jinghao Guo, Xingwei Qu, Yichi Zhang, Chau Yuen

Open Models Built on This Paper

Every model in the SAVRN Model Hub whose card cites this paper, most downloaded first, with what it takes to run each one.

ModelTaskSizeLicenseMonthly downloadsCheapest setup at 16-bit
DN-MOPD-Qwen3.5-2B-baseline-label-160updates
XinLi
Image and text to text 2.2B apache-2.0 — 1x MI300X $1.85/hr
DN-MOPD-Qwen3.5-4B-baseline-label-160updates
XinLi
Image and text to text 4.5B apache-2.0 — 1x MI300X $1.85/hr
DN-MOPD-Qwen3.5-9B-baseline-label-160updates
XinLi
Image and text to text 9.4B apache-2.0 — 1x MI300X $1.85/hr
DN-MOPD-Qwen3.5-2B-160updates
XinLi
Image and text to text 2.2B apache-2.0 — 1x MI300X $1.85/hr
DN-MOPD-Qwen3.5-4B-160updates
XinLi
Image and text to text 4.5B apache-2.0 — 1x MI300X $1.85/hr