paper web signal

Piskorz team: KL direction, not rollouts, drives distillation

TL;DR

  • Rollout policy on its own produced no consistent gain in accuracy, forgetting, or parameter-update sparsity across Llama3 and Qwen2.5 distillation runs.
  • Forward KL stayed robust regardless of who generated the rollouts; reverse KL was substantially more sensitive to the rollout source.
  • Learning rate, not rollout source, governed catastrophic forgetting and update sparsity; on-policy gains on harder Countdown variants did not survive subsequent RLVR.

On-policy rollouts, long cited as the structural advantage of RL fine-tuning over supervised fine-tuning for LLM distillation, do not produce consistent gains once the other knobs are held still.

That is the finding of a controlled study posted to arXiv on 28 September 2026 by Julianna Piskorz, Antonin Berthon and Mihaela van der Schaar. They varied rollout policy, token-level KL divergence direction, and learning rate independently, running strong-to-weak distillation with Llama-3.1-8B teaching Llama-3.2-1B and Qwen2.5-7B teaching Qwen2.5-1.5B-Instruct, on MedReason, Science, Countdown and Numina-MATH.

The reframing is blunt. The paper reports that "token-level KL direction more clearly shapes task performance and output coverage, while learning rate governs forgetting and update sparsity." Forward KL stayed stable regardless of which model generated the rollouts; the authors describe reverse KL as "substantially more sensitive" to that choice. On-policy data did help on the harder Countdown arithmetic variants, but the authors write the advantage "does not reliably persist after subsequent RLVR" training.

The abstract reports no per-benchmark accuracy numbers.