Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

  • 类型:arxiv
  • 标识:2609.20715
  • 链接:https://arxiv.org/abs/2609.20715
  • 主分类:agent
  • 形态:application
  • TLDR:Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets. We ask whether this convention provides the best initialization for subsequent reinforcement learning. We introduce ActObs, which also supervises the observation tokens already present in each trajectory. Although deployed agents never generate observations, learning to predict them encourages the policy to model action consequences without adding data, parameters
  • 副分类:multimodal
  • 待LLM分类:否
  • 标题中文:不要遮蔽环境:观测监督会改变 RL 下智能体的探索行为
  • TLDR中文:Agent 轨迹记录 agent 所执行的动作及其后续结果。然而标准的监督微调(SFT)仅对 agent 生成的动作 token 计算损失,将环境观察作为上下文而非预测目标。本文探讨这一惯例是否为后续强化学习提供了最佳初始化。我们提出 ActObs,对每条轨迹中已存在的观察 token 也进行监督。尽管部署后的 agent 从不生成观察,但学习预测观察 token 能促使策略建模动作后果,且无需新增数据或参数……
  • 来源文件
  • /inbox/tom/_candidates/2026-09-18-agent-rag-longcontext-candidates.json