SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models

  • 类型:arxiv
  • 标识:2608.29974
  • 链接:https://arxiv.org/abs/2608.29974
  • 主分类:multimodal
  • 形态:method
  • TLDR:Detecting hallucinations in Large Vision-Language Models (LVLMs) requires both accurate span localization and well-calibrated confidence scores. Fine-tuned generative VLMs excel at identifying hallucinated text spans but suffer from overconfidence and high inference latency. Discriminative sequence taggers offer deterministic speed and superior calibration but exhibit conservative span recall. We present SpanCalib-VLM, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-
  • 待LLM分类:否
  • 标题中文:SpanCalib-VLM:视觉语言模型中校准的幻觉片段检测
  • TLDR中文:在大视觉语言模型(LVLM)中检测幻觉需要同时具备准确的片段定位和良好校准的置信度分数。微调的生成式 VLM 擅长识别幻觉文本片段,但存在过度自信和推理延迟高的问题。判别式序列标注器具有确定性的速度和更优的校准,但片段召回偏保守。本文提出 SpanCalib-VLM,一种用于 SHROOM-Visions 共享任务的混合双系统,结合了多模态序列标注器(由 XLM-RoBERTa-Large 通过 cross-attention 与 SigLIP 视觉编码器融合而成)……
  • 来源文件
  • /inbox/tom/_candidates/2026-09-02-agent-rag-longcontext-candidates.json