HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

  • 类型:arxiv
  • 标识:2607.18217
  • 链接:https://arxiv.org/abs/2607.18217
  • 主分类:multimodal
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment, and introduces global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens.
  • 待LLM分类:否
  • 标题中文:HOMIE:通过多模态智能增强实现以人-物为中心的视频个性化
  • TLDR中文:HOMIE 提出了一种更优的 MLLM 集成策略,可在不损害文本编码器可控性或引入昂贵重新对齐的前提下,提取参考级关系知识,并在 self-attention 中引入全局多模态引导,使 MLLM 派生的语义特征与 VAE token 更好对齐。
  • 来源文件
  • /inbox/tom/_candidates/2026-07-21-agent-rag-longcontext-candidates.json
  • [S2 enrich]