Can Agents Design Libraries for Agents?

  • 类型:arxiv
  • 标识:2609.36730
  • 链接:https://arxiv.org/abs/2609.36730
  • 主分类:agent
  • 形态:benchmark
  • TLDR:Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. The benchmark spans 242 expert-vali
  • 副分类:evaluation
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-01-agent-rag-longcontext-candidates.json