Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Research

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

arXiv:2609.18708v1 Announce Type: cross Abstract: In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systemati

arXiv cs.AI··Updated just now·34 sightings