τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction

  • 类型:arxiv
  • 标识:2609.04611
  • 链接:https://arxiv.org/abs/2609.04611
  • 主分类:agent
  • 形态:benchmark
  • TLDR:LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and operate internal systems. Notably, the work of building them is increasingly handed to coding agents, yet existing benchmarks say little about whether an AI system can deliver one under the conditions of a real client engagement. We introduce τ^τ-bench (pronounced hyper-tau-bench), a benchmark that makes agent construction the task. A developer agent is given the records a business actually keeps, a client who holds requirements, a production API that operations must run through,
  • 副分类:evaluation
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-07-agent-rag-longcontext-candidates.json