Shanghai AI Lab's SP3O Cures PPO 'Value Flattening' With Sparse Critic Supervision on Three States per Response
Summary
Shanghai AI Laboratory diagnoses 'Value Flattening' in PPO critics, where MC-estimated state values swing sharply while critic predictions stay flat, driven by an implicit variance penalty and redundant temporally-correlated updates. Their SP3O algorithm applies value loss to only three well-separated states per response, showing consistent policy improvements on Qwen3-Base across model sizes and evaluation suites while cutting critic compute.
Originally reported by arxiv.org
Read the original article →Original headline: Shanghai AI Lab's SP3O Cures PPO 'Value Flattening' With Sparse Critic Supervision on Three States per Response