Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
- 类型:arxiv
- 标识:2608.28192
- 链接:https://arxiv.org/abs/2608.28192
- 主分类:multimodal
- 形态:position
- TLDR:Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the target entity throughout that interval. Existing multimodal large language models typically serialize dense localization trajectories autoregressively, causing decoding latency to grow with tube length and allowing localization errors to propagate across time. We introduce Parallel Tube Decoding (PTD), a generative formulation that decomposes grounding into a temporal block followed by time-conditioned spatial blocks decoded simultaneously. This removes both token-level and trajecto
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-08-31-agent-rag-longcontext-candidates.json