PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
- 类型:arxiv
- 标识:2610.00559
- 链接:https://arxiv.org/abs/2610.00559
- 主分类:evaluation
- 形态:benchmark
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:PhysVista is a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human seeing-reasoning-assessment process, highlighting a persistent gap between visual recognition and genuine physical understanding, and pointing toward more principled designs for physically grounded multimodal intelligence.
- 副分类:multimodal
- 待LLM分类:否
- 标题中文:PhysVista:通过感知-推理-评估闭环评测 VLM 的物理智能
- TLDR中文:PhysVista 是一个旨在通过借鉴人类“感知-推理-评估”过程的认知闭环框架来评测 VLM 物理智能的基准,揭示了视觉识别与真实物理理解之间持续存在的差距,并为面向物理基础的多模态智能设计提供了更具原则性的方向。
- 来源文件:
- /inbox/tom/_candidates/2026-10-02-agent-rag-longcontext-candidates.json
- [S2 enrich]