Evo-Bench: Can Language Models Improve Agent Harness?
- 类型:arxiv
- 标识:2608.09096
- 链接:https://arxiv.org/abs/2608.09096
- 主分类:evaluation
- 形态:benchmark
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:Evo-Bench is the first benchmark designed to evaluate models'intrinsic harness-evolving capabilities across Search, Office, and General agent domains, and exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.
- 副分类:agent
- 待LLM分类:否
- 标题中文:Evo-Bench:语言模型能否改进 Agent Harness?
- TLDR中文:Evo-Bench 是首个跨 Search、Office 和 General Agent 领域评估模型内在 harness 演化能力的基准,并暴露了早期饱和等关键时序异常,同时证明所合成的 harness 是高度可迁移的推理结构,能持续提升多样化策略模型。
- 来源文件:
- /inbox/tom/_candidates/2026-08-11-agent-rag-longcontext-candidates.json
- [S2 enrich]