PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

  • 类型:arxiv
  • 标识:2607.24957
  • 链接:https://arxiv.org/abs/2607.24957
  • 主分类:evaluation
  • 形态:benchmark
  • 被引:1
  • 被引来源:Semantic Scholar
  • S2被引:1
  • OpenAlex被引:0
  • 影响力被引:0
  • TLDR:PerceptionBench provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs, by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks and constructing an error taxonomy whose perception branch defines ten atomic perceptual capabilities.
  • OpenAlex ID:W7171653369
  • OpenAlex DOI:10.48550/arxiv.2607.24957
  • DOI:10.48550/arxiv.2607.24957
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://doi.org/10.48550/arxiv.2607.24957
  • OpenAlex更新:2026-08-25
  • 副分类:multimodal
  • 待LLM分类:否
  • 标题中文:PerceptionBench:评估多模态大语言模型的原子级视觉感知
  • TLDR中文:PerceptionBench 通过诊断前沿 MLLMs 在 42 项现有基准上响应中的最早失效点,并构建一个感知分支定义十种原子感知能力的错误分类法,为衡量与诊断 MLLMs 视觉感知边界提供了一项能力级(capability-level)标准。
  • 来源文件
  • /inbox/tom/_candidates/2026-07-29-agent-rag-longcontext-candidates.json
  • [S2 enrich]
  • [OpenAlex backfill]