HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- 类型:arxiv
- 标识:2609.01437
- 链接:https://arxiv.org/abs/2609.01437
- 主分类:evaluation
- 形态:benchmark
- TLDR:As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly termed the agent harness. Changing this harness while holding model weights fixed can substantially alter task performance. Current agent evaluations typically report downstream performance under a chosen harness, leaving a model's ability to develop the harness itself comparatively underexplored. We introduce HarnessDev, a benchmark that shifts the unit of evaluation from task outputs to runnable infrastructure. HarnessDev covers two stages. In
- 副分类:agent
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-03-agent-rag-longcontext-candidates.json