LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows

  • 类型:arxiv
  • 标识:2609.15863
  • 链接:https://arxiv.org/abs/2609.15863
  • 主分类:multimodal
  • 形态:method
  • TLDR:Video diffusion models are stochastic and hard to control: precise content often requires repeated sampling without guaranteed success, and long-horizon scenes drift in appearance, interactions, and temporal coherence. Agentic visual creation provides explicit references, editable 3D scenes, or executable game states for stable control, but does not by itself guarantee high object or character fidelity. Combining the two can enable stable, high-quality generation. To realize this combination, we present LynnReal-Omni, a native multimodal video generation framework built on a 32B shared multimo
  • 副分类:agent
  • 待LLM分类:否
  • 标题中文:LynnReal-Omni:面向Agent视觉工作流的原生多模态视频生成
  • TLDR中文:视频扩散模型具有随机性且难以控制:精确内容往往需要反复采样且无法保证成功,长时场景在外观、交互和时间一致性上会发生漂移。Agent式视觉创作可提供显式参考、可编辑的3D场景或可执行的游戏状态以实现稳定控制,但本身并不能保证高对象或角色保真度。二者结合可实现稳定且高质量的生成。为实现该结合,我们提出 LynnReal-Omni,一个基于32B共享多模态
  • 来源文件
  • /inbox/tom/_candidates/2026-09-15-agent-rag-longcontext-candidates.json