Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

  • 类型:arxiv
  • 标识:2610.01092
  • 链接:https://arxiv.org/abs/2610.01092
  • 主分类:multimodal
  • 形态:benchmark
  • TLDR:Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamically evolve when executing goal-directed actions. While evaluating these capabilities is crucial, existing benchmarks focus mainly on single short actions or step-by-step instructions. This leaves multi-step physical reasoning underexplored, especially in egocentric video generation that requires planning to simulate proper execution to accomplish high-level g
  • 副分类:evaluation
  • 待LLM分类:否
  • 标题中文:Ego2Act:评估自我中心视频生成中的目标导向操作
  • TLDR中文:视频生成模型正越来越多地被探索作为具身规划与学习的世界模拟器。要有效发挥作用,这些模型不仅要生成视觉上吸引人的画面,还要预测在执行目标导向动作时环境的动态演变。虽然评估这些能力至关重要,但现有基准主要聚焦于单一短动作或分步指令,导致多步物理推理仍探索不足,尤其是在需要通过规划来模拟正确执行以完成高层级目标的自我中心视频生成中。
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-03-agent-rag-longcontext-candidates.json