paper web signal

OPD collapse is reward hacking on teacher's implicit signal

TL;DR

  • OPD "improves performance without expanding the student's capabilities" — gains come from easier sampling of correct responses, not new skills.
  • Collapse into long, repetitive output is reward hacking: the teacher amplifies student patterns it rarely generates itself.
  • Masking unhealthy responses during training and SFT initialization each independently mitigate the collapse, the authors report.

On-policy distillation's headline gains and its tendency to collapse into excessively long, repetitive output trace to the same mechanism: the teacher model is functioning as an implicit reward model over student rollouts, not merely as a generator.

That is the argument of a new arXiv paper by Han Cui, Jianhao Yan and co-authors, posted October 2. Their experiments find that OPD "improves performance without expanding the student's capabilities." When the teacher's implicit reward tracks quality, correct responses simply become easier to sample. When it drifts from quality, something else happens. "Reward hacking happens: the implicit reward model amplifies overlong, repetitive student rollouts, even though it rarely generates such text itself," the authors write. The pathological text the teacher almost never produces is exactly what the student is pushed to produce more of.

Two mitigations are reported to work on their own: masking unhealthy responses during training, and SFT initialization. The abstract names no specific base models, teacher models, or benchmark numbers.

The authors close by reframing the problem. OPD "amplifies student behaviors favored by the teacher's implicit feedback, shifting the focus from how well the teacher generates to how reliably it evaluates student rollouts."