ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

  • 类型:arxiv
  • 标识:2609.15635
  • 链接:https://arxiv.org/abs/2609.15635
  • 主分类:multimodal
  • 形态:method
  • TLDR:A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity: MedGemma-27B on 3,199 paired MIMIC-CXR cases from 293 patients, all 14 questions per case (13 finding-specific and one composite), each image replaced by one from another study, usually of the same patient, with question and report fixed. Under an explicit answer instruction, the model's generated answer changes on 4.26 percent of trials with the report and 20.94 per
  • 待LLM分类:否
  • 标题中文:ModaLens:报告条件医学 VLM 中图像敏感性的度量
  • TLDR中文:放射学报告本身已能回答临床问题,因此很难判断视觉语言模型是否也使用了图像。ModaLens 是一种配对图像交换审计,衡量报告可用性如何改变图像敏感性:在293名患者的3,199例配对 MIMIC-CXR 上评估 MedGemma-27B,每例14个问题(13个发现特定问题 + 1个综合问题),每张图像替换为另一研究的图像(通常来自同一患者),问题和报告固定。在显式回答指令下,模型生成答案在有报告时4.26%的试验中发生变化,无报告时为20.94%
  • 来源文件
  • /inbox/tom/_candidates/2026-09-15-agent-rag-longcontext-candidates.json