Modality-Autoregressive World-Action Models

  • 类型:arxiv
  • 标识:2609.17524
  • 链接:https://arxiv.org/abs/2609.17524
  • 主分类:engineering
  • 形态:method
  • TLDR:World-action models (WAMs) jointly model future observations and actions, typically predicting the future as RGB images. Other visual modalities such as depth, pretrained visual features, and point tracks can more efficiently capture geometric, semantic, and motion features. However, how best to combine these modalities within WAMs remains an open question. We introduce ModAR, the first WAM to autoregressively denoise multiple future modalities before predicting actions. This allows each prediction to condition on previously generated modalities. We train from scratch to systematically study h
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-16-agent-rag-longcontext-candidates.json