arXiv:2609.06245 · 评测基准
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
VDiff-Bench:面向细粒度图像差异识别的挑战性 benchmark
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
- 类型:arxiv
- 标识:2609.06245
- 链接:https://arxiv.org/abs/2609.06245
- 主分类:evaluation
- 形态:benchmark
- TLDR:Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often struggle with a basic comparative skill: identifying what has changed between two similar images. We introduce VDiff-Bench, a challenging multiple-choice benchmark for fine-grained Image Difference Identification. VDiff-Bench contains 1,756 four-way questions over image pairs and covers 10 change categories: position, motion, regional image color, overall image color, appearance/disappearance, noise/resolution, texture, substitution/size, OCR/text, a
- 待LLM分类:否
- 标题中文:VDiff-Bench:面向细粒度图像差异识别的挑战性 benchmark
- TLDR中文:多模态大语言模型(MLLMs)在视觉问答等通用视觉理解任务上表现强劲,却在一种基础比较能力上常遇瓶颈:识别两张相似图像之间的差异。我们提出 VDiff-Bench,一个面向细粒度图像差异识别的多选题 benchmark,包含 1,756 道四选一题,覆盖 10 类变化:位置、运动、局部图像颜色、整体图像颜色、出现/消失、噪声/分辨率、纹理、替换/尺寸、OCR/文字……
- 来源文件:
- /inbox/tom/_candidates/2026-09-10-agent-rag-longcontext-candidates.json