OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching
- 类型:arxiv
- 标识:2608.08097
- 链接:https://arxiv.org/abs/2608.08097
- 主分类:llm-infra
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:OasisKV is presented, a memory-centric LLM inference system design that alleviates HBM capacity pressure by decoupling full KV-cache storage from HBM during LLM decoding and observes that future important tokens can be predicted accurately in advance using lookahead tokens drafted by speculative decoding (SD).
- 待LLM分类:否
- 标题中文:OasisKV:通过 Lookahead 稀疏预取将解码端 KV Cache 扩展至 HBM 之外
- TLDR中文:本文提出 OasisKV,一种以显存为中心的 LLM 推理系统设计,通过在 LLM 解码期间将完整 KV-cache 存储与 HBM 解耦来缓解 HBM 容量压力,并观察到未来重要 token 可借助推测解码(SD)所起草的前瞻 token 被提前准确预测。
- 来源文件:
- /inbox/tom/_candidates/2026-08-11-agent-rag-longcontext-candidates.json
- [S2 enrich]