CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation

  • 类型:arxiv
  • 标识:2610.03421
  • 链接:http://arxiv.org/abs/2610.03421v1
  • 主分类:rag
  • 形态:method
  • TLDR:Multimodal large language models (MLLMs) have shown strong visual reasoning abilities, but knowledge-intensive visual question answering often requires external textual evidence beyond the image and the model's parametric knowledge. Existing multimodal RAG systems commonly rely on Top-$K$ retrieval or reranking, which may return redundant passages and provide limited control over whether an answer update is sufficiently supported by the retrieved evidence. We propose \textit{CLIMB}, a training-free inference-time framework for multimodal RAG. CLIMB first constructs a compact complementary evid
  • 副分类:multimodal
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-05-agent-rag-longcontext-candidates.json