Teacher Reward Hacking Explains Both On-Policy Distillation Gains and 99.4% Truncation Collapse

Found first: a primary source the press has not covered yet.

Researchers at Zhejiang University and Westlake University have unified two opposite outcomes of on-policy distillation under a single RL mechanism: the teacher's implicit reward model. In "Gains and Collapse in On-Policy Distillation: A Reinforcement Learning Perspective", Han Cui and colleagues show that whether OPD improves or destroys a student model depends entirely on whether the teacher's implicit preferences align with quality.

What the source says

The team tested student models including DeepSeek-R1-Distill-Qwen-1.5B and Qwen3-1.7B-Base against teachers including JustRL-1.5B, DeepScaleR-1.5B-Preview, and Qwen3-4B, evaluating on AIME24-26. Gains are real but bounded: OPD improved accuracy by up to 16.2 percentage points in the JustRL setting (20.9% to 37.1%), yet a targeted 1,024-response audit found zero problems the student could solve after OPD that it could not already solve, meaning the solvable set did not expand. Collapse is dramatic: with the Qwen3-4B teacher, the student's truncation rate rose from 4.1% to 99.4% and severe repetition jumped from 4.2% to 38.0%, driven by the teacher implicitly rewarding overlong outputs it rarely generates itself. Masking unhealthy rollouts during training recovered an average of 3.08 percentage points of accuracy; SFT warmup initialization pushed average accuracy from 10.16% to 16.38%.

Why it matters

On-policy distillation is now a standard step in LLM post-training pipelines, and its gains and failures have typically been treated as separate phenomena. This paper argues they are the same mechanism, the teacher's implicit reward amplified by RL dynamics, pointing in two directions depending on whether teacher preferences track quality. Practitioners who see repetition and length explosion during OPD are not debugging a separate failure mode; they are observing reward hacking. The two validated fixes, rollout masking and SFT initialization, are straightforward to apply. Code is released at github.com/HancCui/opd_hacking.