Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

  • 类型:arxiv
  • 标识:2609.10355
  • 链接:https://arxiv.org/abs/2609.10355
  • 主分类:multimodal
  • 形态:survey
  • TLDR:Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a computation and memory cost that grows with frame count and context length, limiting deployment in real-time, mobile and resource-constrained settings. This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in paramet
  • 副分类:llm-infra
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-10-agent-rag-longcontext-candidates.json