DN-MOPD Fixes Feedback Imbalance in Qwen3.5 Distillation
TL;DR
- Plain MOPD fails to beat the best single-specialist baseline on Qwen3.5 and captures little of the math specialist's gain, the authors report.
- DN-MOPD keeps the routing but rescales each domain's feedback by its measured spread; on six benchmarks it tops MOPD at every size across three seeds and two length limits.
- Controls show the improvement comes mainly from turning instruction-following feedback down rather than turning mathematics feedback up.
Multi-teacher on-policy distillation does not beat the best single-specialist baseline because the loudest domain hijacks the shared student, according to a new Hugging Face paper on Qwen3.5. The proposed fix, Domain-Normalized MOPD, keeps the routing but rescales each domain's feedback by its measured spread.
In Qwen3.5 experiments at three sizes, plain MOPD 'does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage,' the authors write. Their diagnosis: 'instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates.'
On six public benchmarks, DN-MOPD 'improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain.' Controls with fixed domain weights show the gain 'comes mainly from turning down instruction-following feedback rather than turning up mathematics alone.'
The arxiv version posted September 28. It arrives alongside LSPD in a run of distillation-mechanism papers on our fine-tuning tracker.
Originally reported by huggingface.co
Read the original article →Original headline: HF Paper: Domain-Normalized Multi-Teacher On-Policy Distillation Beats Single-Teacher Baselines