Found first: a primary source the press has not covered yet.
On-policy distillation keeps student models bounded by their RL-trained teacher's capability. Researchers from the University of Science and Technology of China, Rutgers University, Zhejiang University, and one independent contributor propose RIDE (RL-Induced Direction Extrapolation), a representation-space method that exceeds the RL teacher on every one of the four model pairs tested, the only method in the comparison that does so across all four.
What the source says
RIDE computes, layer by layer, the residual between the RL teacher and its pre-RL checkpoint, then regresses the student's hidden states toward targets extrapolated beyond the teacher along that residual direction. The paper tests four base/teacher pairs spanning different scales, architectures, and pre-training lineages: R1-Distill-1.5B, Qwen3-4B, Llama-3.2-3B, and Phi-4-mini. On all four, RIDE's Avg@16 score exceeds the teacher's; margins over OPRD, the strongest non-extrapolation baseline, range from 0.97 to 4.06 points. Output-space extrapolation (ExOPD) collapses on pairs where the teacher is close to its base model because the language-model head attenuates the RL-induced signal anisotropically: the weakest 512 head directions carry 79.8% of the teacher-to-base residual's hidden-state energy but only 30.0% of its centered-logit energy. RIDE avoids this by operating entirely in hidden-state space, where the gradient signal is deterministic.
Why it matters
On-policy distillation has treated the teacher as a fixed target: the student approaches it but cannot pass it. RIDE reframes the teacher as a point on a trajectory, using the pre-RL checkpoint to define a direction of improvement and training the student to continue past the teacher along that direction. That makes distillation a compounding operation rather than a transfer operation. For teams running RL-then-distill pipelines, the result suggests a smaller student can outperform the RL-trained teacher on the tested configurations without a new training run or a stronger teacher. The four pairs cover different architectures, so the effect is not specific to one model family.