Simplified Sparse Attention via Gist Tokens

  • 类型:arxiv
  • 标识:2604.20920
  • 链接:https://arxiv.org/abs/2604.20920
  • 主分类:llm-infra
  • 形态:method
  • 被引:2
  • 被引来源:Semantic Scholar
  • S2被引:2
  • OpenAlex被引:0
  • 影响力被引:0
  • TLDR:Simplified Sparse Attention is introduced, a simpler approach to sparse attention that requires no architectural changes that outperforms full attention in retrieval-augmented generation and extends to a hierarchical gist-of-gist variant that achieves log-linear decoding complexity while maintaining or improving accuracy at high compression ratios up to 32x.
  • OpenAlex ID:W7155526994
  • OpenAlex DOI:10.48550/arxiv.2604.20920
  • DOI:10.48550/arxiv.2604.20920
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://doi.org/10.48550/arxiv.2604.20920
  • OpenAlex更新:2026-07-19
  • 待LLM分类:否
  • 标题中文:基于 Gist Token 的简化稀疏注意力
  • TLDR中文:提出 Simplified Sparse Attention——一种更简洁的稀疏注意力方案,无需任何架构改动,在检索增强生成中优于全注意力;并扩展为分层 gist-of-gist 变体,在高达 32× 的高压缩比下保持或提升精度,同时实现对数级解码复杂度。
  • 来源文件
  • /inbox/tom/_candidates/2026-06-30-rag-retrieval-reranking-candidates.json
  • [S2 enrich]
  • [OpenAlex backfill]