Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World

  • 类型:arxiv
  • 标识:2609.23038
  • 链接:https://arxiv.org/abs/2609.23038
  • 主分类:multimodal
  • 形态:method
  • TLDR:Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transitions caused by object motion and viewpoint changes and integrate them over long trajectories to maintain an updated spatial state, yet existing VLMs remain limited in both capabilities. Current spatial training primarily focuses on static questions about object attributes and spatial relations, providing limited direct supervision for state transitions; in contrast, interaction trajectories naturally connect
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-24-agent-rag-longcontext-candidates.json