FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

  • 类型:arxiv
  • 标识:2608.19758
  • 链接:http://arxiv.org/abs/2608.19758v1
  • 主分类:llm-infra
  • 形态:application
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:This paper introduces a mean correction term that effectively suppresses the approximation error, keeping performance degradation manageable even at extreme sparsity levels, and redesigns the sparse attention operator with PackGQA memory access, warp specialization, and pingpong pipelining.
  • 待LLM分类:否
  • 标题中文:FlashPrefill V2:面向长上下文 LLM 服务的块稀疏 Prefill 注意力
  • TLDR中文:本文引入一个均值修正项,有效抑制近似误差,即使在极端稀疏度下也能将性能下降保持在可控范围;并使用 PackGQA 内存访问、warp specialization 和 pingpong 流水线重新设计稀疏注意力算子。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-21-agent-rag-longcontext-candidates.json
  • [S2 enrich]