LVMT: Video Mask Transformer for Long-term Video Segmentation

  • 类型:arxiv
  • 标识:2609.34895
  • 链接:https://arxiv.org/abs/2609.34895
  • 主分类:multimodal
  • 形态:method
  • TLDR:Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this limitation is caused by (i) the inability of their temporal propagation mechanism to adaptively select the object information that is propagated across time, and (ii) their inability to be trained on long videos due to memory requirements and vanishing gradients. To address the first limitation, we propose to use a lightweight GRU-based temporal propagation module that can learn to select which information it keeps in memory and propagates across time
  • 待LLM分类:否
  • 标题中文:LVMT:用于长期视频分割的视频掩码 Transformer
  • TLDR中文:现有在线视频分割方法难以在具有长期遮挡的冗长复杂视频中稳定跟踪目标。我们将这一局限归因于:(i)其时间传播机制无法自适应选择跨时间传播的目标信息;(ii)受内存需求及梯度消失影响,难以在长视频上训练。针对第一个局限,我们提出使用一个轻量级的基于 GRU 的时间传播模块,使其能够学习选择保留在记忆中并跨时间传播的信息
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-05-agent-rag-longcontext-candidates.json