arXiv:2609.37690 · 多模态
Honeycomb: Constant-Size Scene Memory Representation for Video World Models
Honeycomb:面向视频世界模型的恒定大小场景记忆表征
Honeycomb: Constant-Size Scene Memory Representation for Video World Models
- 类型:arxiv
- 标识:2609.37690
- 链接:https://arxiv.org/abs/2609.37690
- 主分类:multimodal
- 形态:method
- TLDR:Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memories accumulate RGB observations or latent features, increasing storage requirements as generation proceeds. We introduce Honeycomb, a video world model built on HexMemory, our proposed low-rank representation for storing scene features in a fixed-size memory with a total of six spatial and spatiotemporal planes. A feed-forward writer maps each generated chunk into new plane features. As the spatial coverage or temporal range expands, we warp the previous planes
- 待LLM分类:否
- 标题中文:Honeycomb:面向视频世界模型的恒定大小场景记忆表征
- TLDR中文:视频世界模型需要持久的场景记忆以在长时域视频生成中维持一致性。现有空间记忆累积 RGB 观测或 latent 特征,随着生成进行存储需求不断增长。我们提出 Honeycomb,一种基于 HexMemory 构建的视频世界模型——HexMemory 是我们提出的低秩表征,用于在固定大小的记忆中以总共六个空间与时空平面存储场景特征。一个前馈 writer 将每个生成块映射为新的平面特征。随着空间覆盖范围或时间范围的扩展,我们对先前平面进行 warp……
- 来源文件:
- /inbox/tom/_candidates/2026-10-03-agent-rag-longcontext-candidates.json