OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching

  • 类型:arxiv
  • 标识:2608.08097
  • 链接:https://arxiv.org/abs/2608.08097
  • 主分类:llm-infra
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:OasisKV is presented, a memory-centric LLM inference system design that alleviates HBM capacity pressure by decoupling full KV-cache storage from HBM during LLM decoding and observes that future important tokens can be predicted accurately in advance using lookahead tokens drafted by speculative decoding (SD).
  • 待LLM分类:否
  • 标题中文:OasisKV:通过 Lookahead 稀疏预取将解码端 KV Cache 扩展至 HBM 之外
  • TLDR中文:本文提出 OasisKV,一种以显存为中心的 LLM 推理系统设计,通过在 LLM 解码期间将完整 KV-cache 存储与 HBM 解耦来缓解 HBM 容量压力,并观察到未来重要 token 可借助推测解码(SD)所起草的前瞻 token 被提前准确预测。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-11-agent-rag-longcontext-candidates.json
  • [S2 enrich]