Papers · organized/paper_cards

论文

1 张论文卡片 · 多模态 · 评测集

开放获取 全部 绿色 · 724
Multimodal 补充候选
arXiv:2508.17398 多模态 评测集 被引 3 · S2

本文提出 DashboardQA,这是首个明确设计用于评估视觉-语言 GUI Agent 对真实世界仪表板理解与交互能力的基准,结果表明交互式仪表板推理对所有受评估的 VLM 而言都是一项具有挑战性的任务。DashboardQA is introduced, the first benchmark explicitly designed to assess how vision-language GUI agents comprehend and interact with real-world dashboards, and indicates that interactive dashboard reasoning is a challenging task overall for all the VLMs evaluated.