HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement
- 类型:arxiv
- 标识:2607.18217
- 链接:https://arxiv.org/abs/2607.18217
- 主分类:multimodal
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment, and introduces global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens.
- 待LLM分类:否
- 标题中文:HOMIE:通过多模态智能增强实现以人-物为中心的视频个性化
- TLDR中文:HOMIE 提出了一种更优的 MLLM 集成策略,可在不损害文本编码器可控性或引入昂贵重新对齐的前提下,提取参考级关系知识,并在 self-attention 中引入全局多模态引导,使 MLLM 派生的语义特征与 VAE token 更好对齐。
- 来源文件:
- /inbox/tom/_candidates/2026-07-21-agent-rag-longcontext-candidates.json
- [S2 enrich]