paper web signal

SLCA-GRPO fixes tool-call credit leak, +9.15pp on τ²-Bench

TL;DR

  • Standard GRPO broadcasts one trajectory-level advantage to every token, mixing rewards meant for tool calls with rewards meant for user-facing summaries.
  • Segment-Locked Credit Assignment splits the advantage at segment boundaries and routes execution credit to tool tokens, preference credit to summary tokens.
  • On a 7B backbone the fix adds 2.53 pp on Toucan, 1.36 pp on BFCL, and 9.15 pp on τ²-Bench versus vanilla GRPO.

A large slice of the reward signal in GRPO tool-calling training is being routed to the wrong tokens, according to a new paper on arxiv. On a 7B backbone, fixing the routing lifts τ²-Bench by 9.15 points versus standard GRPO.

The problem: tool-calling agents produce mixed outputs. Some tokens are structured API invocations, others are natural-language replies to the user. GRPO gives all of them the same trajectory-level advantage. The paper describes standard algorithms as ones that "indiscriminately broadcast a homogeneous trajectory-level scalar advantage to all tokens," causing "cross-segment credit misattribution and brittle optimization."

The authors' fix, Segment-Locked Credit Assignment, splits the advantage at segment boundaries and pairs it with a hierarchical reward scheme. It "routes execution advantages to tool tokens and preference advantages to summary tokens." On the 7B model, they report gains of 2.53 pp on the in-domain Toucan set, 1.36 pp on the Berkeley Function-Calling Leaderboard, and the 9.15 pp τ²-Bench jump.

To avoid burning real API calls during training, the authors also built a Schema-Guided LLM Simulator (SGLS) to stand in for tool environments. The retrieved abstract does not report results at larger model sizes, and does not break out how much of the τ²-Bench gain comes from segment-locking on its own versus the hierarchical reward.