PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop

  • 类型:arxiv
  • 标识:2610.00559
  • 链接:https://arxiv.org/abs/2610.00559
  • 主分类:evaluation
  • 形态:benchmark
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:PhysVista is a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human seeing-reasoning-assessment process, highlighting a persistent gap between visual recognition and genuine physical understanding, and pointing toward more principled designs for physically grounded multimodal intelligence.
  • 副分类:multimodal
  • 待LLM分类:否
  • 标题中文:PhysVista:通过感知-推理-评估闭环评测 VLM 的物理智能
  • TLDR中文:PhysVista 是一个旨在通过借鉴人类“感知-推理-评估”过程的认知闭环框架来评测 VLM 物理智能的基准,揭示了视觉识别与真实物理理解之间持续存在的差距,并为面向物理基础的多模态智能设计提供了更具原则性的方向。
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-02-agent-rag-longcontext-candidates.json
  • [S2 enrich]