All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation

  • 类型:arxiv
  • 标识:2609.27901
  • 链接:https://arxiv.org/abs/2609.27901
  • 主分类:multimodal
  • 形态:method
  • TLDR:Video is a rich representation of a physical event, capturing appearance, geometry, motion, and temporal evolution. Other modalities, such as 3D body motion or audio, encode narrower aspects of the same event. We find that joint multimodal diffusion transformers exhibit a corresponding asymmetry in cross-modal correspondence: companion modalities develop strong correspondences to video, but the reciprocal correspondences through which they constrain video remain substantially weaker. We express both directions as comparable correspondence distributions over video tokens and define their disagr
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-24-agent-rag-longcontext-candidates.json