arxiv.org web signal

Shanghai AI Lab's SP3O Cures PPO 'Value Flattening' With Sparse Critic Supervision on Three States per Response

Summary

Shanghai AI Laboratory diagnoses 'Value Flattening' in PPO critics, where MC-estimated state values swing sharply while critic predictions stay flat, driven by an implicit variance penalty and redundant temporally-correlated updates. Their SP3O algorithm applies value loss to only three well-separated states per response, showing consistent policy improvements on Qwen3-Base across model sizes and evaluation suites while cutting critic compute.