GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?

  • 类型:arxiv
  • 标识:2608.21833
  • 链接:https://arxiv.org/abs/2608.21833
  • 主分类:agent
  • 形态:benchmark
  • 被引:4
  • 被引来源:Semantic Scholar
  • S2被引:4
  • OpenAlex被引:0
  • 影响力被引:0
  • TLDR:GameXpert-Bench is introduced, which operationalizes the three lifecycle stages of game development with a coding agent as three complementary benchmark tracks, and finds current agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, and preserving functionality across changes.
  • OpenAlex ID:W7204193705
  • OpenAlex DOI:10.48550/arxiv.2608.21833
  • DOI:10.48550/arxiv.2608.21833
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://arxiv.org/pdf/2608.21833
  • OpenAlex更新:2026-08-31
  • 副分类:evaluation
  • 待LLM分类:否
  • 标题中文:GameXpert-Bench:编码 Agent 距专家级游戏开发还有多远?
  • TLDR中文:本文提出 GameXpert-Bench,将游戏开发的三个生命周期阶段以 coding agent 操作化为三条互补的基准轨道,并发现现有 agent 在生成可玩基础框架和实现显式需求方面更可靠,而在发现缺陷、验证运行时行为以及跨变更保持功能一致性方面能力较弱。
  • 来源文件:
  • /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json
  • [S2 enrich]
  • [OpenAlex backfill]