ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

  • 类型:arxiv
  • 标识:2609.10895
  • 链接:https://arxiv.org/abs/2609.10895
  • 主分类:multimodal
  • 形态:benchmark
  • TLDR:Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as the decision coreof household robots. Existing evaluations, however, probe intuitive physics passively through question answering over videos, or target deliberate, long-horizon tasks such as navigation and rearrangement; none measure whether a model can turn physical understanding into immediate, safety-critical action. We introduce ReactHuman, the first physics-grounded benc
  • 副分类:evaluation
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-15-rag-retrieval-reranking-candidates.json
  • /inbox/tom/_candidates/2026-09-15-agent-rag-longcontext-candidates.json