PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training
- 类型:arxiv
- 标识:2610.04616
- 链接:https://arxiv.org/abs/2610.04616
- 主分类:multimodal
- 形态:method
- TLDR:A vision--language--action (VLA) policy can complete complex tasks while ignoring the evidence that should determine its actions. An object held near the wrist camera can displace the instructed target. Language and action show the same pattern: a familiar noun can trigger the operation it was paired with in training even after the verb changes, and a gripper that closed on nothing may lift anyway. We call these dependencies modality shortcuts: regularities in successful demonstrations make visual, lexical, or motor cues sufficient to predict expert actions without the task evidence needed for
- 副分类:engineering
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-10-06-agent-rag-longcontext-candidates.json