Found first: a primary source the press has not covered yet.
Researchers from Zhejiang University, Westlake University, the Shanghai Innovation Institute, and Shanghai AI Laboratory have found that GRPO-trained reasoning models reinforce tokens that shift under trivial prompt rewrites alongside tokens that reflect genuine reasoning. Their method, SCAPO, improves Qwen3-4B-Base AIME 2024-2026 accuracy by 5.63 percentage points over GRPO, as described in the paper posted September 30, 2026.
What the source says
The team applied semifactual prompt interventions: the same problem, the same correct answer, surface features rephrased. They measured how much each token's probability drifts under those rewrites. As a diagnostic, they sampled one response per prompt from a frozen Qwen3-4B-Base on 1,000 questions drawn from DAPO-Math-17K, then suppressed high-drift token candidates during decoding without updating model weights. Accuracy rose from 15.8% to 30.0%. SCAPO bakes this signal into training: it computes token-level drift scores and reduces the GRPO advantage for unstable tokens during early training, without granting extra credit for stability alone. On AIME 2024-2026, Qwen3-4B-Base reaches 27.71% under SCAPO versus 22.08% under GRPO; Qwen3-1.7B-Base reaches 11.88% versus 7.71%. SCAPO achieves the best results on most evaluated mathematics benchmarks and on all evaluated out-of-distribution benchmarks among the compared methods.
Why it matters
Standard GRPO assigns the same outcome-derived advantage to every token in a response, so tokens that co-occur with correct answers get reinforced regardless of whether they contributed to the reasoning. The semifactual framing gives a causal measure of this: if a token's probability swings under a rewrite that should not change anything, the token is tracking surface noise. The inference-only diagnostic makes the problem concrete: nearly doubling accuracy without touching weights shows the brittleness is already baked into what the model learned, not a decoding artifact. The gains carrying to out-of-distribution benchmarks suggests the method improves generalization rather than fitting the training distribution more tightly.