Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

  • 类型:arxiv
  • 标识:2609.32577
  • 链接:https://arxiv.org/abs/2609.32577
  • 主分类:agent
  • 形态:method
  • TLDR:Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task requirements. This leaves the policy without a learning signal that favors clean, targeted implementations over those containing unnecessary or out-of-scope changes. We introduce GAGAR, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic sampling
  • 待LLM分类:否
  • 标题中文:面向代码 Agent 强化学习的分组智能评分与优势再分配
  • TLDR中文:代码 Agent 的强化学习(RL)常使用可执行测试提供二元奖励。在此奖励下,GRPO 将同一 rollout 组内通过测试的轨迹赋予相同优势,忽略了实现质量与任务要求遵守程度的差异,导致策略缺少偏好简洁、针对性实现而非含不必要或越界修改的学习信号。我们提出 GAGAR,一个面向代码 Agent RL 的质量感知信用再分配框架,基于动态采样
  • 来源文件:
  • /inbox/tom/_candidates/2026-09-29-agent-rag-longcontext-candidates.json