SLCA-GRPO fixes tool-call credit leak, +9.15pp on τ²-Bench
TL;DR
- Standard GRPO broadcasts one trajectory-level advantage to every token, mixing rewards meant for tool calls with rewards meant for user-facing summaries.
- Segment-Locked Credit Assignment splits the advantage at segment boundaries and routes execution credit to tool tokens, preference credit to summary tokens.
- On a 7B backbone the fix adds 2.53 pp on Toucan, 1.36 pp on BFCL, and 9.15 pp on τ²-Bench versus vanilla GRPO.
A large slice of the reward signal in GRPO tool-calling training is being routed to the wrong tokens, according to a new paper on arxiv. On a 7B backbone, fixing the routing lifts τ²-Bench by 9.15 points versus standard GRPO.
The problem: tool-calling agents produce mixed outputs. Some tokens are structured API invocations, others are natural-language replies to the user. GRPO gives all of them the same trajectory-level advantage. The paper describes standard algorithms as ones that "indiscriminately broadcast a homogeneous trajectory-level scalar advantage to all tokens," causing "cross-segment credit misattribution and brittle optimization."
The authors' fix, Segment-Locked Credit Assignment, splits the advantage at segment boundaries and pairs it with a hierarchical reward scheme. It "routes execution advantages to tool tokens and preference advantages to summary tokens." On the 7B model, they report gains of 2.53 pp on the in-domain Toucan set, 1.36 pp on the Berkeley Function-Calling Leaderboard, and the 9.15 pp τ²-Bench jump.
To avoid burning real API calls during training, the authors also built a Schema-Guided LLM Simulator (SGLS) to stand in for tool environments. The retrieved abstract does not report results at larger model sizes, and does not break out how much of the τ²-Bench gain comes from segment-locking on its own versus the hierarchical reward.
Originally reported by paper
Read the original article →Original headline: PKU Patches Hidden Credit-Leak Bug in Tool-Calling RL: +9.15pp on τ²-Bench at 7B