HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management

  • 类型:arxiv
  • 标识:2608.07009
  • 链接:http://arxiv.org/abs/2608.07009v1
  • 主分类:llm-infra
  • 形态:position
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:HiSparse is merged into upstream SGLang and evaluated across three sparse-attention families (DSA, NSA, and Quest) on H200, B200, and GH200 platforms: it improves peak generation throughput by up to 4.7x on long-context workloads while preserving comparable per-token latency and reducing time-to-first-token at high load.
  • 待LLM分类:否
  • 标题中文:HiSparse:通过分层 KV 缓存管理扩展稀疏注意力解码
  • TLDR中文:HiSparse 已合入上游 SGLang,并在 H200、B200 与 GH200 平台上针对 DSA、NSA、Quest 三类稀疏注意力家族进行评估:在长上下文负载下峰值生成吞吐提升最高达 4.7×,同时保持可比的 per-token 延迟,并降低高负载下的 time-to-first-token。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-10-agent-rag-longcontext-candidates.json
  • [S2 enrich]