SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

  • 类型:arxiv
  • 标识:2609.08149
  • 链接:https://arxiv.org/abs/2609.08149
  • 主分类:evaluation
  • 形态:benchmark
  • TLDR:SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents' true coding ability. We present SWE-Bench Pro Verified, a verified version of SWE-Bench Pro that addresses both problems.
  • 副分类:agent
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-10-agent-rag-longcontext-candidates.json