TLDR
可验证奖励的强化学习 (RLVR) 主要通过交互后的标量结果奖励将 Agent 经验转化为学习信号。然而对于组相对目标,当所有推演获得相同奖励时,该信号即消失,尽管这些轨迹可能包含关于任务内容及 Agent 失败方式的有用信息。我们提出一个互补问题:事后回顾能否教会 Agent 在行动前本可预见的内容?我们引入前瞻学习,利用事后经验从行动前的Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-