VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
- 类型:arxiv
- 标识:2609.04355
- 链接:https://arxiv.org/abs/2609.04355
- 主分类:multimodal
- 形态:method
- TLDR:Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demanding precision and repeatability. Applying real-world online reinforcement learning (RL) to VLA post-training enables autonomous trial-and-error improvement beyond demonstrations alone, but exposes two bottlenecks: 1) unreliable value signals can induce policy drift; 2) large-VLA overhead constrains throughput and sample efficiency. To address these challenges, we present VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorith
- 副分类:engineering
- 待LLM分类:否
- 标题中文:[标题中文] VLA-Precision:面向视觉-语言-动作模型高效真实世界在线强化学习的非对称协同自举
- TLDR中文:预训练的 vision-language-action (VLA) 模型支持广泛的操控任务,但在需要精确性与可重复性的任务中仍不可靠。将真实世界在线强化学习 (RL) 应用于 VLA 后训练,可实现超越单纯示教的自主试错改进,但也暴露出两个瓶颈:1) 不可靠的价值信号会导致策略漂移;2) 大型 VLA 的开销限制了吞吐量和样本效率。为应对这些挑战,我们提出了 VLA-Precision,一个高效的真实世界在线 RL 框架,其核心是 Asymmetric Co-Bootstrapping (ACoB) 算法……
- 来源文件:
- /inbox/tom/_candidates/2026-09-28-agent-rag-longcontext-candidates.json