Found first: a primary source the press has not covered yet.
A controlled study testing Llama3 and Qwen2.5 on scientific, medical, and arithmetic reasoning tasks finds no consistent accuracy advantage from on-policy rollout generation in knowledge distillation. The real driver, according to a new paper from Piskorz, Berthon, and van der Schaar, is the direction of the KL divergence objective.
What the source says
The study isolates rollout policy from KL objective across Llama3 and Qwen2.5 model families, testing on scientific, medical, and arithmetic reasoning benchmarks including Countdown. Forward KL is robust to rollout policy, performing stably whether data comes from the student or teacher model. Reverse KL is substantially more sensitive and favors student-generated rollouts. On-policy data does improve generalization to harder Countdown variants under both KL directions, but this advantage does not reliably persist after subsequent RLVR training. Learning rate is the primary driver of catastrophic forgetting and parameter-update sparsity.
Why it matters
On-policy rollout generation is expensive and widely cited as a key advantage of RL-based fine-tuning over supervised methods. This study finds that advantage is not consistent and that KL objective direction is a more consequential design choice. For practitioners building distillation pipelines, this shifts the focus: forward KL offers stability across rollout conditions; reverse KL's sensitivity means rollout policy is a meaningful variable specifically for that objective. The finding that learning rate governs forgetting has direct implications for continual training and multi-stage fine-tuning recipes.