ExplainBench: Evaluating Code Explanations from Agents

  • 类型:arxiv
  • 标识:2607.26451
  • 链接:https://arxiv.org/abs/2607.26451
  • 主分类:agent
  • 形态:survey
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:This work proposes ExplainBench, a benchmark to automatically evaluate explanations from coding agents, based on the intuition that informative explanations should enable an LLM to correctly answer questions, allowing quantitative comparison of explanation quality between agents.
  • 副分类:evaluation
  • 待LLM分类:否
  • 标题中文:ExplainBench:评估 Agent 的代码解释
  • TLDR中文:提出 ExplainBench,一个自动评估 coding agent 解释的基准,基于信息性解释应能让 LLM 正确回答问题的直觉,实现 agent 之间解释质量的量化比较。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-05-agent-rag-longcontext-candidates.json
  • [S2 enrich]