Agent Plasticity: Measuring Self-Improvement Through Experience

  • 类型:arxiv
  • 标识:2610.08902
  • 链接:https://arxiv.org/abs/2610.08902
  • 主分类:agent
  • 形态:method
  • TLDR:AI agents increasingly operate in environments where they can diagnose failures and improve through experience, yet existing evaluations largely measure what an agent can do at a fixed point in time rather than how effectively it learns. Evaluating self-improvement requires answering three questions: does future performance improve and generalize beyond the interactions that enabled learning; how efficiently are new capabilities acquired; and where does the self-improvement process break down? To answer these questions, we study self-improvement in a controlled setting where agents amortize pa
  • 副分类:evaluation
  • 待LLM分类:否
  • 标题中文:Agent 可塑性:通过经验衡量自我改进
  • TLDR中文:AI Agent 日益在能够诊断失败并通过经验改进的环境中运行,然而现有评估主要衡量 Agent 在固定时间点能做什么,而非其学习效果如何。评估自我改进需要回答三个问题:未来性能是否改进并泛化至学习交互之外;新能力获取效率如何;自我改进过程在哪里失效?为回答这些问题,我们在受控环境中研究自我改进,其中 Agent 分摊
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-09-agent-rag-longcontext-candidates.json