ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding

  • 类型:arxiv
  • 标识:2609.02780
  • 链接:https://arxiv.org/abs/2609.02780
  • 主分类:multimodal
  • 形态:application
  • TLDR:Streaming video understanding is a critical capability for real-world applications, including embodied intelligence, autonomous driving, industrial monitoring, surveillance and early warning, and wearable assistants. However, processing continuous video streams with multimodal large language models (MLLMs) is computationally expensive. Existing efforts have explored reducing streaming overhead through visual token pruning, token merging, quantization, on-demand frame retrieval, and context offloading. However, most existing methods overlook the dimension of model depth. Repeatedly executing fu
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-07-agent-rag-longcontext-candidates.json