WorldGuide: Goal-Directed Video World Model for Procedural Task Execution

  • 类型:arxiv
  • 标识:2610.12459
  • 链接:https://arxiv.org/abs/2610.12459
  • 主分类:multimodal
  • 形态:method
  • TLDR:Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Open-loop generation cannot adapt to execution outcomes, while existing closed-loop systems often rely on pretrained executors or indirect verification. This leaves a gap between deciding an action and successfully realizing it. We formulate procedural video generation as closed-loo
  • 待LLM分类:否
  • 标题中文:WorldGuide:面向过程化任务执行的目标导向视频世界模型
  • TLDR中文:视频生成器和基于视频的世界模型能够合成合理的视觉轨迹,但长时程过程化任务要求生成过程能适应已实际生成的内容。模型必须根据其生成的状态决定下一步动作、执行该动作,并识别任务何时完成。开环(Open-loop)生成无法适应执行结果,而现有闭环(closed-loop)系统往往依赖预训练执行器或间接验证,导致动作决策与成功执行之间存在鸿沟。我们将过程化视频生成形式化为闭环
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-09-agent-rag-longcontext-candidates.json