MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

  • 类型:arxiv
  • 标识:2609.38078
  • 链接:https://arxiv.org/abs/2609.38078
  • 主分类:multimodal
  • 形态:method
  • TLDR:Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost. This motivates us to ask: Can a general-purpose VLM itself operate a robot mo
  • 副分类:agent
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-05-agent-rag-longcontext-candidates.json