ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering

  • 类型:arxiv
  • 标识:2609.02486
  • 链接:http://arxiv.org/abs/2609.02486v1
  • 主分类:rag
  • 形态:method
  • TLDR:Document Visual Question Answering (DocVQA) often leverages Retrieval-Augmented Generation (RAG), where late-interaction encoders are commonly used to identify document pages relevant to a user query, before answer generation by a Large Vision-Language Model (LVLM). Existing approaches typically retrieve a fixed top-$k$ number of pages regardless of query complexity, which increases LVLM latency and may degrade answer accuracy. We introduce ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-$k$ retrieval method for late-interaction visual document retrieval. ViSAR operates
  • 副分类:engineering
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-03-agent-rag-longcontext-candidates.json