CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation
- 类型:arxiv
- 标识:2610.03421
- 链接:http://arxiv.org/abs/2610.03421v1
- 主分类:rag
- 形态:method
- TLDR:Multimodal large language models (MLLMs) have shown strong visual reasoning abilities, but knowledge-intensive visual question answering often requires external textual evidence beyond the image and the model's parametric knowledge. Existing multimodal RAG systems commonly rely on Top-$K$ retrieval or reranking, which may return redundant passages and provide limited control over whether an answer update is sufficiently supported by the retrieved evidence. We propose \textit{CLIMB}, a training-free inference-time framework for multimodal RAG. CLIMB first constructs a compact complementary evid
- 副分类:multimodal
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-10-05-agent-rag-longcontext-candidates.json