Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
- 类型:arxiv
- 标识:2609.18708
- 链接:https://arxiv.org/abs/2609.18708
- 主分类:engineering
- 形态:position
- TLDR:In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated from multiple Monte Carlo continuations, change sharply across intermediate states while critic predictions remain comparatively flat. We further observe this phenomenon in a controlled FrozenLake environment and find that it becomes more pronounced as the state space grows. Our theoretical and empiric
- 待LLM分类:是
- 来源文件:
- /inbox/tom/_candidates/2026-09-17-agent-rag-longcontext-candidates.json