Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
- 类型:arxiv
- 标识:2607.07608
- 链接:https://arxiv.org/abs/2607.07608
- 主分类:multimodal
- 形态:method
- 被引:1
- 被引来源:Semantic Scholar
- S2被引:1
- OpenAlex被引:0
- 影响力被引:0
- TLDR:LaMem-VLA is introduced, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning, and enables memory to directly participate in VLA reasoning and guide action generation under a bounded context.
- OpenAlex ID:W7167852388
- OpenAlex DOI:10.48550/arxiv.2607.07608
- DOI:10.48550/arxiv.2607.07608
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://doi.org/10.48550/arxiv.2607.07608
- OpenAlex更新:2026-07-19
- 副分类:rag
- 待LLM分类:否
- 标题中文:面向机器人操作的视觉-语言-动作模型中的双潜在记忆
- TLDR中文:提出 LaMem-VLA,一种以潜在记忆为核心框架的方法,将历史经验重建为潜在记忆 token,并直接与 VLA 推理交织,使记忆能够在有界上下文下直接参与 VLA 推理并引导动作生成。
- 来源文件:
- /inbox/tom/_candidates/2026-07-09-agent-rag-longcontext-candidates.json
- [S2 enrich]
- [OpenAlex backfill]