ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
- 类型:arxiv
- 标识:2609.18487
- 链接:https://arxiv.org/abs/2609.18487
- 主分类:multimodal
- 形态:position
- TLDR:Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-09-17-agent-rag-longcontext-candidates.json