Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models
- 类型:arxiv
- 标识:2608.23478
- 链接:https://arxiv.org/abs/2608.23478
- 主分类:multimodal
- 形态:method
- TLDR:Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which motor command was demonstrated while leaving implicit the local objective served by the behavior under the instruction. Future-based supervision enriches action learning with frames, latent observations, trajectories, or motion representations, but these signals capture particular realizations of what may happen rather than the shared semantic objective of the forthcoming behavior. We propose Intention Distillation (I
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-08-31-agent-rag-longcontext-candidates.json