CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering
- 类型:arxiv
- 标识:2609.32600
- 链接:https://arxiv.org/abs/2609.32600
- 主分类:agent
- 形态:benchmark
- TLDR:Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studied in isolation, leaving this integrated development process underexplored. Diagnosing a runtime interaction failure requires agents to connect visual observations with the responsible code, then use the application again to verify the repair. We introduce CUA-SWE, a benchmark, environment, and eval
- 副分类:evaluation
- 待LLM分类:否
- 标题中文:CUA-SWE:当 Computer-Use Agents 遇见可视化软件工程
- TLDR中文:软件开发不仅仅是编辑代码:开发者反复运行软件、与界面交互、可视化检查其行为,并依据这些观察决定下一步修改内容以及修改是否有效。现有 coding agents 和 computer-use agents 很大程度上被独立研究,这一集成开发过程尚未被充分探索。诊断运行时交互故障要求 agents 将可视化观察与对应代码关联起来,然后再次使用应用程序来验证修复。我们提出 CUA-SWE,一个 benchmark、环境与评估框架
- 来源文件:
- /inbox/tom/_candidates/2026-10-01-agent-rag-longcontext-candidates.json