LSPD Frames Distillation as RL, Matches OPD on 25% Rollouts
TL;DR
- LSPD frames on-policy distillation as a reinforcement-learning problem and proves a sharp Õ(log K) regret bound under online exploration.
- The method reports average gains of +1.59 points in Avg@16 across six mathematical reasoning benchmarks.
- A fully off-policy variant matches vanilla OPD performance using only the first 25% of rollout batches.
On-policy distillation, viewed through a reinforcement-learning lens, admits a sharp Õ(log K) regret bound and a fully off-policy variant that matches vanilla OPD using only the first 25% of rollout batches, according to a preprint by Shangzhe Li and six co-authors posted to arXiv on September 28.
The framework, called Least Square Policy Distillation, incorporates 'optimistic exploration and off-policy trajectory reuse' around the standard distillation objective. The authors report LSPD 'outperforms existing distillation methods' across six mathematical reasoning benchmarks, with 'average gains of +1.59 points in Avg@16,' and stronger Pass@k results as k grows toward 64.
The 25% rollout figure is the eye-catch. If it holds up outside the math suite, teams burning teacher-inference cycles for distillation have a lever to cut those cycles by three-quarters. The six benchmarks listed are all mathematical reasoning; the abstract mentions no evaluations on code, agents, or tool use. Code is on GitHub. It lands two days after Hugging Face's domain-normalized multi-teacher on-policy distillation paper, one of fifty fine-tuning stories we've tracked in the last ninety days.
Originally reported by arxiv.org
Read the original article →Original headline: LSPD Paper Reframes On-Policy Distillation as RL, Matches OPD on 25% of Rollouts