SANTA++: Sampling Attention through Representative Keys
- 类型:arxiv
- 标识:2609.35629
- 链接:http://arxiv.org/abs/2609.35629v1
- 主分类:engineering
- 形态:method
- TLDR:Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection without scanning the entire key-value (KV) cache. Cached keys are organized into teams, and the query scores one representative from each team to decide which teams to sample. We compute exact attention scores within the sampled teams and reweight each team's contribution by the inverse of its inclusio
- 待LLM分类:否
- 标题中文:SANTA++: 通过代表性 Key 进行采样的注意力机制
- TLDR中文:注意力往往集中在上下文中一小部分 token 上,但每个 query 关注的关键子集各不相同。为利用这种动态结构,我们提出 SANTA++,一种免训练的随机注意力方法,通过代表性 key 进行内存高效的选择,无需扫描整个 KV cache。缓存的 key 被组织成若干 team,query 对每个 team 中的代表性 key 打分以决定采样哪些 team。我们在采样得到的 team 内计算精确的注意力分数,并通过其采样概率的倒数对各 team 的贡献进行重新加权。
- 来源文件:
- /inbox/tom/_candidates/2026-09-30-agent-rag-longcontext-candidates.json