paper web signal

RIDE lets students match or beat their RL-trained teachers

TL;DR

  • RIDE extrapolates in representation space, computing per-layer residuals between the teacher and its pre-RL checkpoint rather than operating on output logits.
  • Across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the teacher on every pair.
  • Output-space extrapolation, the alternative generalized variant, degrades the student whenever the teacher is close to its base, per the paper.

A distillation method called RIDE pushes a student past the mean score of its RL-trained teacher on every tested pair, by moving the training signal from output space into representation space. RIDE stands for RL-Induced Direction Extrapolation.

The paper, posted to arXiv by Hao Li and collaborators, describes the mechanism this way: "at every layer and token position, RIDE computes the residual between the teacher and its pre-RL checkpoint and regresses the student's hidden states toward targets displaced beyond the teacher along this residual." Reinforcement learning, the authors argue, shifts a model's internal representations relative to its base checkpoint, and that shift can be measured at every layer. RIDE treats the shift as a direction the student can keep moving along, not a fixed target to match.

The empirical claim is narrow but clean. Across "four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages," the method "approaches or exceeds the RL-trained teacher on every pair and is the only method whose mean does so." Output-space extrapolation, the alternative generalized variant, "degrades the student whenever the teacher is close to its base," per the paper. The abstract names no specific pairs, publishes no per-model numbers, and does not disclose the extrapolation coefficient.