Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
- 类型:arxiv
- 标识:2607.26326
- 链接:https://arxiv.org/abs/2607.26326
- 主分类:multimodal
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:For the coarse attributes the authors study, MLLMs encode the visual evidence but cannot reliably control their reliance on it, indicating that for the coarse attributes they study, MLLMs cannot reliably control their reliance on it.
- 副分类:engineering
- 待LLM分类:否
- 标题中文:标题 -> 标题中文:看见还是知道?多模态大语言模型中的视觉上下文敏感性
- TLDR中文:在作者研究的粗粒度属性上,MLLM 编码了视觉证据但无法可靠控制对其的依赖
- 来源文件:
- /inbox/tom/_candidates/2026-08-05-agent-rag-longcontext-candidates.json
- [S2 enrich]