EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?

  • 类型:arxiv
  • 标识:2609.04280
  • 链接:https://arxiv.org/abs/2609.04280
  • 主分类:evaluation
  • 形态:benchmark
  • TLDR:Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EVOHARNESSBENCH, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents). Unlike existing continual-learning benchmarks for agents, which typically place non-stationarity (i.e., what changes over time) in the task stream while keeping the harness fixed, EVOHARNESSBENCH places non-stationarity in the
  • 副分类:agent
  • 待LLM分类:否
  • 标题中文:EVOHARNESSBENCH:你的 Agent 能否跟上不断演进的 Harness?
  • TLDR中文:现代基于 LLM 的 Agent 通过由工具、可复用 skill 和专家 Agent 组成的 Harness 来运作,该 Harness 决定了它们能观察到什么、能做什么。在实践中,随着新能力的不断加入,这一 Harness 持续演进。我们提出 EVOHARNESSBENCH,一个用于在 Harness 受控演化条件下评估 Agent 的基准,覆盖三个维度(工具、skill 和 Agent)。与现有面向 Agent 的持续学习基准不同——后者通常将非平稳性(即随时间变化的部分)放在任务流中而保持 Harness 固定——EVOHARNESSBENCH 将非平稳性放在
  • 来源文件
  • /inbox/tom/_candidates/2026-09-10-agent-rag-longcontext-candidates.json