WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
- 类型:arxiv
- 标识:2608.04964
- 链接:https://arxiv.org/abs/2608.04964
- 主分类:multimodal
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:WorldCycle is introduced, a self-verifiable RL framework that constructs closed action cycles and their repeated executions from ordinary action sequences, and optimizes two complementary rewards: a spatial closure reward enforcing symmetry between mirrored forward and reverse segments, and a temporal consistency reward aligning states across repeated cycle executions.
- 待LLM分类:否
- 标题中文:WorldCycle:面向长视野视频世界模型的自可验证强化学习
- TLDR中文:提出 WorldCycle,一种自验证 RL 框架,从普通动作序列中构建闭合动作循环及其重复执行,并优化两个互补奖励:空间闭合奖励(强制镜像的前向与反向片段之间的对称性)以及时间一致性奖励(对齐多次循环执行间的状态)。
- 来源文件:
- /inbox/tom/_candidates/2026-08-06-agent-rag-longcontext-candidates.json
- [S2 enrich]