BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

  • 类型:arxiv
  • 标识:2608.05042
  • 链接:https://arxiv.org/abs/2608.05042
  • 主分类:multimodal
  • 形态:application
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:BridgeVLA++ is developed by equipping BridgeVLA with a unified spatio-temporal memory architecture that models persistent spatial context and temporal interaction history that can reason over observation histories while preserving BridgeVLA's data efficiency and generalization capabilities.
  • 待LLM分类:否
  • 标题中文:BridgeVLA++:用于 3D 操作的、数据高效、可泛化且具备记忆增强的 Vision-Language-Action 框架
  • TLDR中文:在 BridgeVLA 基础上开发 BridgeVLA++,引入统一的时空记忆架构,建模持久化的空间上下文与时间交互历史,使其可在保留 BridgeVLA 数据效率与泛化能力的同时对观测历史进行推理。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-06-agent-rag-longcontext-candidates.json
  • [S2 enrich]