Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

  • 类型:arxiv
  • 标识:2609.03820
  • 链接:https://arxiv.org/abs/2609.03820
  • 主分类:multimodal
  • 形态:method
  • TLDR:Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system keeps only a small fixed slice of that pool. Which frames survive that slice is usually treated as a preprocessing detail; we test whether it should be. Published selectors make the comparison hard because they change the frame scorer, the prompt boundary, the resolution policy, and the answering model all at once. We hold each fixed and vary one decision at a time: selection, spatial compression, and reinvestment of the savings, across six training-free selection rules, three l
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-04-agent-rag-longcontext-candidates.json