It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
- 类型:arxiv
- 标识:2609.37863
- 链接:https://arxiv.org/abs/2609.37863
- 主分类:multimodal
- 形态:method
- TLDR:Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all. The guidelines require the label to be decided from the sentence alone, so no image should change any answer. We expected each image to pull a
- 待LLM分类:否
- 标题中文:关键不在图像本身:无关上下文会扰动 VLM 评判且不提供有效信息
- TLDR中文:Vision-language models (VLMs) 日益被用于替代人类标注者,这使得可替代性测试应反映模型本身而非偶然的评估条件。我们提出 MIST(Misleading-Image Stress Test):包含 200 个英文句子,每个围绕一个可作比喻或字面解读的短语构建,并配以一张展示其对应解读的对齐图像、一张展示相反解读的误导图像,或不配图。标注指南要求标签仅依据句子本身决定,因此任何图像都不应改变答案。我们预期每张图像都会
- 来源文件:
- /inbox/tom/_candidates/2026-10-01-agent-rag-longcontext-candidates.json