Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
arXiv:2609.18708v1 Announce Type: cross Abstract: In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systemati
arXiv cs.AI··Updated just now·34 sightings