GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?
- 类型:arxiv
- 标识:2608.21833
- 链接:https://arxiv.org/abs/2608.21833
- 主分类:agent
- 形态:benchmark
- TLDR:Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especially demanding because program logic, visual and audio content, interfaces, interaction and playability must function together in one executable artifact. Measuring this capability therefore requires evaluation of both game product and the development process. Existing benchmarks often assess the game development capabilities of LLMs by evaluating the final artifact or an isolated development stage. Our analysis of complete human-agent development
- 副分类:evaluation
- 待LLM分类:否
- 标题中文:GameXpert-Bench:编码 Agent 距专家级游戏开发还有多远?
- TLDR中文:近期大语言模型(LLM)已能作为编码 Agent,根据自然语言请求构建完整游戏。游戏开发尤为严苛,因为程序逻辑、视觉与音频内容、界面、交互和可玩性必须在同一可执行制品中协同工作。因此衡量该能力需要同时评测游戏产品与开发过程。现有基准通常通过评估最终制品或孤立的开发阶段来评测 LLM 的游戏开发能力。我们对完整人机协作开发过程的分析
- 来源文件:
- /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json