Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
- 类型:arxiv
- 标识:2609.32577
- 链接:https://arxiv.org/abs/2609.32577
- 主分类:agent
- 形态:method
- TLDR:Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task requirements. This leaves the policy without a learning signal that favors clean, targeted implementations over those containing unnecessary or out-of-scope changes. We introduce GAGAR, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic sampling
- 待LLM分类:否
- 标题中文:面向代码 Agent 强化学习的分组智能评分与优势再分配
- TLDR中文:代码 Agent 的强化学习(RL)常使用可执行测试提供二元奖励。在此奖励下,GRPO 将同一 rollout 组内通过测试的轨迹赋予相同优势,忽略了实现质量与任务要求遵守程度的差异,导致策略缺少偏好简洁、针对性实现而非含不必要或越界修改的学习信号。我们提出 GAGAR,一个面向代码 Agent RL 的质量感知信用再分配框架,基于动态采样
- 来源文件:
- /inbox/tom/_candidates/2026-09-29-agent-rag-longcontext-candidates.json