TLDR
AI Agent 日益在能够诊断失败并通过经验改进的环境中运行,然而现有评估主要衡量 Agent 在固定时间点能做什么,而非其学习效果如何。评估自我改进需要回答三个问题:未来性能是否改进并泛化至学习交互之外;新能力获取效率如何;自我改进过程在哪里失效?为回答这些问题,我们在受控环境中研究自我改进,其中 Agent 分摊AI agents increasingly operate in environments where they can diagnose failures and improve through experience, yet existing evaluations largely measure what an agent can do at a fixed point in time rather than how effectively it learns. Evaluating self-improvement requires answering three questions: does future performance improve and generalize beyond the interactions that enabled learning; how efficiently are new capabilities acquired; and where does the self-improvement process break down? To answer these questions, we study self-improvement in a controlled setting where agents amortize pa