CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering

  • 类型:arxiv
  • 标识:2609.32600
  • 链接:https://arxiv.org/abs/2609.32600
  • 主分类:agent
  • 形态:benchmark
  • TLDR:Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studied in isolation, leaving this integrated development process underexplored. Diagnosing a runtime interaction failure requires agents to connect visual observations with the responsible code, then use the application again to verify the repair. We introduce CUA-SWE, a benchmark, environment, and eval
  • 副分类:evaluation
  • 待LLM分类:否
  • 标题中文:CUA-SWE:当 Computer-Use Agents 遇见可视化软件工程
  • TLDR中文:软件开发不仅仅是编辑代码:开发者反复运行软件、与界面交互、可视化检查其行为,并依据这些观察决定下一步修改内容以及修改是否有效。现有 coding agents 和 computer-use agents 很大程度上被独立研究,这一集成开发过程尚未被充分探索。诊断运行时交互故障要求 agents 将可视化观察与对应代码关联起来,然后再次使用应用程序来验证修复。我们提出 CUA-SWE,一个 benchmark、环境与评估框架
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-01-agent-rag-longcontext-candidates.json