EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents

  • 类型:arxiv
  • 标识:2609.01281
  • 链接:https://arxiv.org/abs/2609.01281
  • 主分类:multimodal
  • 形态:application
  • TLDR:Vision-language-action (VLA) models map visual observations and language instructions directly to robot actions, but long-horizon tasks require more than action prediction. An agent must coordinate perception, planning, execution, progress verification, and recovery as the physical state evolves. An action prediction or a model-generated skill decision does not, by itself, guarantee that the proposed operation is valid in the current state or that its outcome will be verified. We propose EmbodiedSkills, a unified framework that treats each skill decision as an execution proposal: the runtime c
  • 副分类:agent
  • 待LLM分类:否
  • 标题中文:EmbodiedSkills:用于编排、训练和部署 VLA Agent 的统一框架
  • TLDR中文:视觉-语言-动作(VLA)模型将视觉观测与语言指令直接映射为机器人动作,但长时任务需要的远不止动作预测。Agent 须随物理状态演变协调感知、规划、执行、进度验证与恢复。动作预测或模型生成的 skill 决策本身并不能保证所提操作在当前状态下有效,也不能保证其结果会被验证。我们提出 EmbodiedSkills,一个将每次 skill 决策视为执行提案的统一框架:运行时持续……
  • 来源文件
  • /inbox/tom/_candidates/2026-09-08-agent-rag-longcontext-candidates.json