WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

  • 类型:arxiv
  • 标识:2608.04964
  • 链接:https://arxiv.org/abs/2608.04964
  • 主分类:multimodal
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:WorldCycle is introduced, a self-verifiable RL framework that constructs closed action cycles and their repeated executions from ordinary action sequences, and optimizes two complementary rewards: a spatial closure reward enforcing symmetry between mirrored forward and reverse segments, and a temporal consistency reward aligning states across repeated cycle executions.
  • 待LLM分类:否
  • 标题中文:WorldCycle:面向长视野视频世界模型的自可验证强化学习
  • TLDR中文:提出 WorldCycle,一种自验证 RL 框架,从普通动作序列中构建闭合动作循环及其重复执行,并优化两个互补奖励:空间闭合奖励(强制镜像的前向与反向片段之间的对称性)以及时间一致性奖励(对齐多次循环执行间的状态)。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-06-agent-rag-longcontext-candidates.json
  • [S2 enrich]