AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

  • 类型:arxiv
  • 标识:2608.06729
  • 链接:https://arxiv.org/abs/2608.06729
  • 主分类:multimodal
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:AtlasVLA is a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state and decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.
  • 待LLM分类:否
  • 标题中文:AtlasVLA:面向视觉—语言—动作模型的持久世界—自我状态建模
  • TLDR中文:AtlasVLA 是一种新框架,通过持久化的世界-自我状态从直接反应式操作转向主动推理,显著优于多视角基线,在 LIBERO-Long 上取得 9.4% 的绝对成功率提升,在真实世界长周期任务中取得 17.5% 的提升。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-13-agent-rag-longcontext-candidates.json
  • [S2 enrich]