Piskorz team: KL direction, not rollouts, drives distillation
TL;DR
- Rollout policy on its own produced no consistent gain in accuracy, forgetting, or parameter-update sparsity across Llama3 and Qwen2.5 distillation runs.
- Forward KL stayed robust regardless of who generated the rollouts; reverse KL was substantially more sensitive to the rollout source.
- Learning rate, not rollout source, governed catastrophic forgetting and update sparsity; on-policy gains on harder Countdown variants did not survive subsequent RLVR.
On-policy rollouts, long cited as the structural advantage of RL fine-tuning over supervised fine-tuning for LLM distillation, do not produce consistent gains once the other knobs are held still.
That is the finding of a controlled study posted to arXiv on 28 September 2026 by Julianna Piskorz, Antonin Berthon and Mihaela van der Schaar. They varied rollout policy, token-level KL divergence direction, and learning rate independently, running strong-to-weak distillation with Llama-3.1-8B teaching Llama-3.2-1B and Qwen2.5-7B teaching Qwen2.5-1.5B-Instruct, on MedReason, Science, Countdown and Numina-MATH.
The reframing is blunt. The paper reports that "token-level KL direction more clearly shapes task performance and output coverage, while learning rate governs forgetting and update sparsity." Forward KL stayed stable regardless of which model generated the rollouts; the authors describe reverse KL as "substantially more sensitive" to that choice. On-policy data did help on the harder Countdown arithmetic variants, but the authors write the advantage "does not reliably persist after subsequent RLVR" training.
The abstract reports no per-benchmark accuracy numbers.
Originally reported by paper
Read the original article →Original headline: Controlled Study: KL Direction—Not On-Policy Rollouts—Drives LLM Distillation Gains