HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management
- 类型:arxiv
- 标识:2608.07009
- 链接:http://arxiv.org/abs/2608.07009v1
- 主分类:llm-infra
- 形态:position
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:HiSparse is merged into upstream SGLang and evaluated across three sparse-attention families (DSA, NSA, and Quest) on H200, B200, and GH200 platforms: it improves peak generation throughput by up to 4.7x on long-context workloads while preserving comparable per-token latency and reducing time-to-first-token at high load.
- 待LLM分类:否
- 标题中文:HiSparse:通过分层 KV 缓存管理扩展稀疏注意力解码
- TLDR中文:HiSparse 已合入上游 SGLang,并在 H200、B200 与 GH200 平台上针对 DSA、NSA、Quest 三类稀疏注意力家族进行评估:在长上下文负载下峰值生成吞吐提升最高达 4.7×,同时保持可比的 per-token 延迟,并降低高负载下的 time-to-first-token。
- 来源文件:
- /inbox/tom/_candidates/2026-08-10-agent-rag-longcontext-candidates.json
- [S2 enrich]