CorrGRPO fixes GRPO normalization bias in multi-reward RL
TL;DR
- CorrGRPO swaps the raw pairwise covariance terms in GRPO's normalizer for Pearson correlation coefficients, so large-scale rewards no longer swallow smaller ones.
- On LeetCodeDataset coding with Qwen2.5-Coder, the paper reports CorrGRPO gaining 2.09 to 4.21 Pass@1 points over GRPO across 0.5B to 7B models.
- The method is evaluated across code generation, tool calling, and agent security, with code released at HKUST-KnowComp/CorrGRPO.
GRPO normalizes a batch of rollouts for the same prompt by dividing each rollout's centered reward by the within-group standard deviation. When rewards are combined, say correctness plus a style score, that denominator becomes the sum of all pairwise covariances, and whichever reward carries the largest scale ends up dominating it. In a new arxiv paper, Wenbin Hu and co-authors argue that is a design flaw rather than a feature: "correlated rewards with large scales can dominate this normalization and suppress signals from smaller-scale rewards."
Their fix, called CorrGRPO, swaps those raw covariance terms for Pearson correlation coefficients. The centered total reward stays the same; what moves is how the denominator weighs each component. The method, the authors write, lets "advantage magnitudes to adapt to reward correlations without the normalization being dominated by large-scale reward components."
The paper evaluates on three multi-reward domains, code generation, tool calling, and agent security, with models from 0.5B to 8B parameters. On LeetCodeDataset coding runs with Qwen2.5-Coder, CorrGRPO gains 2.09 to 4.21 Pass@1 points over the GRPO baseline across model scales, with the 7B rising from 47.28% to 51.49%. The tool-calling and agent-security evaluations show similar directional lifts, including a jump from 31.38% to 47.88% joint accuracy on Agent Security Bench for Qwen2.5-7B.
The specific gains don't appear in the abstract itself; they live in the paper's experiment tables, and the authors release code at HKUST-KnowComp/CorrGRPO.
Originally reported by paper
Read the original article →Original headline: CorrGRPO Fixes Multi-Reward Normalization in GRPO, Gains +4.2 Pass@1 on 7B Coding