CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs

  • 类型:arxiv
  • 标识:2609.08345
  • 链接:https://arxiv.org/abs/2609.08345
  • 主分类:engineering
  • 形态:method
  • TLDR:Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produces thousands of redundant visual tokens whose cost grows with every view. Existing visual token pruners fall into two families, each limited in the 3D multi-view setting. Learned importance methods rank tokens by attention or encoder features; because redundancy here is fundamentally spatial, they keep near-duplicate tokens from a few prominent regions and leave most of the scene unrepresented. Voxelization methods im
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-09-agent-rag-longcontext-candidates.json