EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

  • 类型:arxiv
  • 标识:2608.05565
  • 链接:https://arxiv.org/abs/2608.05565
  • 主分类:multimodal
  • 形态:position
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:EffectLearner is proposed, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser that achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.
  • 待LLM分类:否
  • 标题中文:EffectLearner:面向真实世界视频物体移除的世界感知物体-效果推理
  • TLDR中文:EffectLearner 是一个语义推理增强框架,结合基于 VLM 的 Object-Effect Reasoner 与基于 DiT 的 Video Eraser,在 EffectWorld-Eval 和具有挑战性的 EffectWorld-Wild 上均取得明显优势,证明其能在复杂真实场景中实现高质量的视频物体擦除。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-07-agent-rag-longcontext-candidates.json
  • [S2 enrich]