arXiv:2610.10845 · LLM 基础设施
Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute
AI 的真实长期记忆:比重计算更快更便宜的 5000 万 token 窗口
Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute
- 类型:arxiv
- 标识:2610.10845
- 链接:https://arxiv.org/abs/2610.10845
- 主分类:llm-infra
- 形态:method
- TLDR:A large language model can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent. We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it. We ran it on 50,000,000 tokens of real public text, served through vLLM on one NVIDIA H100, with Gemma 4 12B and Gemma 4 31B. Every block we probed was loaded back from the encrypted store with no recompute (100 of 100,
- 待LLM分类:否
- 标题中文:AI 的真实长期记忆:比重计算更快更便宜的 5000 万 token 窗口
- TLDR中文:大语言模型只能使用其上下文窗口中容纳的文本,并且每次发送 prompt 时都会重新计算其内部的 key-value (KV) 状态。我们测试了一个记忆层,即公开包 galahad-kv,它将每个约 16000 token 块的 KV 状态保存到加密的本地 NVMe 磁盘,并在之后按字节精确、无重计算地加载回来。我们在 5000 万 token 的真实公共文本上运行,通过单块 NVIDIA H100 上的 vLLM 提供服务,配合 Gemma 4 12B 和 Gemma 4 31B。我们探测的每个块都从加密存储中无重计算地加载回来(100 个中的 100 个
- 来源文件:
- /inbox/tom/_candidates/2026-10-09-agent-rag-longcontext-candidates.json