AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling

  • 类型:arxiv
  • 标识:2608.26623
  • 链接:https://arxiv.org/abs/2608.26623
  • 主分类:evaluation
  • 形态:benchmark
  • TLDR:LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically study LLM-as-a-judge reliability for agentic tool-calling over workflow DAGs, as distinct from the broader LLM-as-a-judge task of open-ended text or preference evaluation. The benchmark comprises 3,808 instances spanning six DAG topologies and three difficulty tiers, evaluated with five generators (3B-70B open-weight models and GPT-5.4) and six judges (20B to frontier s
  • 副分类:agent
  • 待LLM分类:否
  • 标题中文:AgentJudgeBench:一个用于评估 LLM 法官在智能体工具调用方面表现的多难度基准
  • TLDR中文:LLM 法官被广泛用于评估智能体工具调用系统,但它们在结构化、依赖驱动的流程上的可靠性仍未得到充分研究。我们提出 AgentJudgeBench,这是首个系统性地研究 LLM-as-a-judge 在工作流 DAG 上进行智能体工具调用可靠性的基准,与开放式文本或偏好评估这一更广泛的 LLM-as-a-judge 任务有所区别。该基准包含 3,808 个实例,涵盖六种 DAG 拓扑和三个难度等级,使用五个生成器(3B-70B 开源权重模型以及 GPT-5.4)和六个法官(20B 到前沿模型
  • 来源文件
  • /inbox/tom/_candidates/2026-09-03-agent-rag-longcontext-candidates.json