WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
- 类型:arxiv
- 标识:2610.12459
- 链接:https://arxiv.org/abs/2610.12459
- 主分类:multimodal
- 形态:method
- TLDR:Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Open-loop generation cannot adapt to execution outcomes, while existing closed-loop systems often rely on pretrained executors or indirect verification. This leaves a gap between deciding an action and successfully realizing it. We formulate procedural video generation as closed-loo
- 待LLM分类:否
- 标题中文:WorldGuide:面向过程化任务执行的目标导向视频世界模型
- TLDR中文:视频生成器和基于视频的世界模型能够合成合理的视觉轨迹,但长时程过程化任务要求生成过程能适应已实际生成的内容。模型必须根据其生成的状态决定下一步动作、执行该动作,并识别任务何时完成。开环(Open-loop)生成无法适应执行结果,而现有闭环(closed-loop)系统往往依赖预训练执行器或间接验证,导致动作决策与成功执行之间存在鸿沟。我们将过程化视频生成形式化为闭环
- 来源文件:
- /inbox/tom/_candidates/2026-10-09-agent-rag-longcontext-candidates.json