τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
- 类型:arxiv
- 标识:2608.16885
- 链接:https://arxiv.org/abs/2608.16885
- 主分类:multimodal
- 形态:method
- TLDR:Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce τ_0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution
- 副分类:llm-infra
- 待LLM分类:否
- 标题中文:τ_0-VLA:基于世界模型引导测试时计算的分层机器人基础模型
- TLDR中文:长周期机器人操控要求机器人既能可靠执行各项技能,又能在一系列长时间任务中将其连贯编排。多数分层视觉-语言-动作(VLA)模型仅通过单次前向过程做出每个决策,缺乏将额外算力分配给困难或关键抉择的机制。我们提出 τ_0-VLA,一种分层机器人基础模型,将高层子任务生成建模为可通过世界模型引导的测试时计算来扩展算力的推理问题。在每次推理时,高层策略借助执行……
- 来源文件:
- /inbox/tom/_candidates/2026-08-22-agent-rag-longcontext-candidates.json
- /inbox/tom/_candidates/2026-08-23-agent-rag-longcontext-candidates.json