paper web signal

SCAPO Boosts Qwen3-4B AIME Accuracy by 5.63pp Over GRPO

TL;DR

  • SCAPO beat GRPO on AIME 2024-2026 by 5.63 percentage points on Qwen3-4B-Base and 4.17 points on Qwen3-1.7B-Base.
  • The paper traces the gap to GRPO assigning the same outcome-derived advantage to every response token, reinforcing prompt-sensitive habits alongside real reasoning.
  • SCAPO measures token probability drift under prompt rewrites that preserve the problem, then shrinks credit for unstable tokens during early RL training.

Reinforcement learning with verifiable rewards has improved the reasoning of large language models, yet their 'predictions remain sensitive to task-irrelevant prompt features.' That is the opening claim of a new preprint from Junshu Pan and seven co-authors, dated 30 September 2026.

The authors probe the sensitivity using 'semifactual prompt interventions that preserve the underlying problem and its answer,' then measure how a fixed response's token probabilities drift under those rewrites. 'Our analysis reveals substantial variation in token-level sensitivity,' they write, and they report that 'suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights.'

That cuts against Group Relative Policy Optimization, the dominant RLVR algorithm. GRPO, the paper argues, 'assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning.' Their fix, Semifactual Credit-Augmented Policy Optimization, uses the drift measurements as normalized stability scores to shrink the advantage given to the shakiest tokens during early training, while 'granting no additional credit for stability alone.'

On Qwen3-4B-Base, SCAPO improves AIME 2024-2026 accuracy over GRPO by 5.63 percentage points; on Qwen3-1.7B-Base, by 4.17. SCAPO 'achieves the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among the compared methods' at both scales. The abstract reports no tests on larger base models or on non-Qwen families.