MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation
- 类型:arxiv
- 标识:2609.38078
- 链接:https://arxiv.org/abs/2609.38078
- 主分类:multimodal
- 形态:method
- TLDR:Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost. This motivates us to ask: Can a general-purpose VLM itself operate a robot mo
- 副分类:agent
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-10-05-agent-rag-longcontext-candidates.json