Cliff Paper: Zeroing In on the First Mistake in a Rollout Boosts RL-Trained LLMs 15% Over Distillation
Summary
Cliff is a reward-shaping method for RL with verifiable rewards that pinpoints where reasoning first breaks and assigns positive token-level advantages to the correct prefix, negative feedback afterward. Reported gains: +15% over distillation, +7% over standard outcome-based reward on LLM reasoning benchmarks. Authored by Peixuan Han et al.
Originally reported by huggingface.co
Read the original article →Original headline: Cliff Paper: Zeroing In on the First Mistake in a Rollout Boosts RL-Trained LLMs 15% Over Distillation