The Attention Triangle in Audio-Video Models

  • 类型:arxiv
  • 标识:2609.03586
  • 链接:https://arxiv.org/abs/2609.03586
  • 主分类:multimodal
  • 形态:method
  • TLDR:Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This e
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-07-agent-rag-longcontext-candidates.json