arXiv:2609.15800 · RAG 检索增强
Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation
驾驭稀疏证据:通过显式上下文选择与整合的 Agentic Visual RAG
Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation
- 类型:arxiv
- 标识:2609.15800
- 链接:http://arxiv.org/abs/2609.15800v1
- 主分类:rag
- 形态:method
- TLDR:Visual Retrieval-Augmented Generation (VRAG) empowers models to navigate and answer queries about visually rich documents by retrieving relevant page images as visual evidence and reasoning over their content. However, effectively utilizing this visual evidence is usually impeded by two main challenges. First, answer-relevant evidence is sparse and may be concentrated in a small region of one page or dispersed across multiple pages. Second, existing agentic methods often generate answers based on raw exploration trajectories or compressed textual memories rather than an explicitly organized se
- 副分类:agent
- 待LLM分类:否
- 标题中文:驾驭稀疏证据:通过显式上下文选择与整合的 Agentic Visual RAG
- TLDR中文:Visual RAG 使模型能够通过检索相关页面图像作为视觉证据并对其内容进行推理,从而浏览并回答关于视觉丰富文档的查询。然而,有效利用这些视觉证据通常受两大挑战阻碍:其一,与答案相关的证据稀疏,可能集中在一页的小区域内,也可能分散于多页之中;其二,现有 agentic 方法通常基于原始探索轨迹或压缩的文本记忆生成答案,而非基于显式组织的证...
- 来源文件:
- /inbox/tom/_candidates/2026-09-16-agent-rag-longcontext-candidates.json