NVIDIA/kvpress · 上手攻略
- 仓库:
NVIDIA/kvpress - 链接:https://github.com/NVIDIA/kvpress
- 分类:AI · LLM 推理优化 / KV Cache 压缩
- 作者:Tom
- 更新:2026-10-02
§0 速览
KVPress 是 NVIDIA 出品的 KV Cache 压缩工具包,通过在 prefilling 阶段对注意力层的 Key-Value 缓存做选择性剪枝,把大上下文模型的显存占用压到原来的几分之一。全程无需训练、即装即用,支持 14+ 种压缩算法统一评测 pipeline。核心价值:让 128K~1M token 的长上下文推理在消费级 / 单卡服务器上变得可能。
§1 是什么 / 解决什么问题
1.1 KV Cache 的困境
Transformer 自回归生成时,每生成一个新 token,都需要把所有历史 token 的 Key(K)和 Value(V)重新参与注意力计算。KV Cache 把这些 K/V 结果缓存起来避免重复计算,代价是 显存随上下文长度线性增长:
Size(KV) = 2 × precision × n_layers × n_heads × d × n_tokens
以 Llama 3.1-70B(bfloat16)处理 1M token 为例:
Size(KV) = 2 × 2 × 80 × 8 × 128 × 1,000,000 ≈ 327 GB
模型权重本身 140 GB,加上 327 GB 的 KV Cache = 需要约 470 GB 显存——单卡 A100(80 GB)根本装不下。这堵墙被称为 KV Cache Memory Wall。
1.2 KVPress 的解法
KVPress 在 prefilling(前向传播)阶段 通过可插拔的 "Press"(压缩器)分析每一层 attention 的重要性,剪掉最不重要的 KV 对,把 cache 压缩到指定比例。例如 compression_ratio=0.5 表示把 KV cache 压缩到原来的一半。
官方 benchmark(Llama 3.1-8B,128K 上下文):
| 配置 | 峰值显存 |
|---|---|
| 无压缩 | ~45 GB |
| KVPress 50% 压缩 | ~37 GB |
| 压缩后解码加速 | ✅ 是(cache 更小,解码更快) |
另一个关键场景:AIME 数学竞赛推理(KV budget 仅 2,048 tokens)下,ExpectedAttentionPress 在压缩比相同的情况下精度明显优于 SnapKV 等竞品。
§2 快速安装
2.1 pip 安装(推荐生产使用)
pip install kvpress
2.2 本地源码安装(含全部可选依赖)
git clone https://github.com/NVIDIA/kvpress.git
cd kvpress
uv sync --extra eval --extra flash-attn
--extra eval 安装评测依赖,--extra flash-attn 安装 FlashAttention-2 支持。
⚠️ 要求:CUDA 显卡(推荐 A100/H100 或同等算力),Python 3.10+,Transformers 库。
§3 核心用法
3.1 Pipeline 方式(最简)
KVPress 注册了一个名为 kv-press-text-generation 的自定义 transformers pipeline,最容易上手:
from transformers import pipeline
from kvpress import ExpectedAttentionPress
model = "Qwen/Qwen3-8B"
pipe = pipeline(
"kv-press-text-generation",
model=model,
device_map="auto",
dtype="auto"
)
context = "A very long text you want to compress once and for all"
question = "\nA question about the compressed context" # 可选
press = ExpectedAttentionPress(compression_ratio=0.5)
answer = pipe(context, question=question, press=press)["answer"]
上例中,compression 只作用于 context tokens,不影响 question 部分,用户可以同一个压缩结果问多个不同问题。
3.2 Decoding 压缩(实验性)
KVPress 还支持在生成(decoding)阶段周期性压缩缓存,降低长输出的显存峰值:
from transformers import pipeline
from kvpress import KnormPress, DecodingPress
device = "cuda:0"
model = "meta-llama/Llama-3.1-8B-Instruct"
model_kwargs = {"attn_implementation": "flash_attention_2"}
pipe = pipeline(
"kv-press-text-generation",
model=model,
device=device,
model_kwargs=model_kwargs
)
decoding_press = DecodingPress(
base_press=KnormPress(),
compression_interval=10, # 每 10 步压缩一次
target_size=512 # 压缩后 cache 大小(token 数)
)
context = "A very long text you want to compress during generation"
question = "Tell me a long story about this context"
response = pipe(context, question=question, press=decoding_press)["answer"]
⚠️ 注意:DecodingPress 是实验性功能,只支持 ScorerPress(如 KnormPress)作为 base press,不支持 StreamingLLMPress 等非打分类 press。
3.3 底层 Hook 方式(高级 / 自定义 press)
每个 press 通过 forward hook 注册到 attention 层,用作 context manager 即可:
from kvpress import KnormPress
press = KnormPress()
with press(model): # 注册所有 attention 层的 hook
output = model.generate(**inputs)
# 退出 with 块时 hook 自动解除
§4 支持的 Press 一览
所有 press 均无需训练,继承自 BasePress;多数基于重要性打分的 press 继承自 ScorerPress。
| Press | 类型 | 核心原理 | 论文 |
|---|---|---|---|
| RandomPress | Scorer | 随机打分(baseline) | — |
| KnormPress | Scorer | key 向量的逆范数 | arXiv:2406.11430 |
| SnapKVPress | Scorer | last query 的平均注意力权重 | arXiv:2404.14469 |
| ExpectedAttentionPress | Scorer | 估计生成阶段期望注意力 | arXiv:2510.00636v1 |
| StreamingLLMPress | Base | 只保留初始 token + 最近 token | arXiv:2309.17453 |
| TOVAPress | Scorer | last query 在各头的注意力权重 | arXiv:2401.06104 |
| ObservedAttentionPress | Scorer | prefilling 阶段观察到的平均注意力 | arXiv:2306.14048 |
| QFilterPress | Scorer | key 投影到 query 的 SVD 主分量 | arXiv:2503.02812 |
| PyramidKVPress | Scorer | 底层多 cache、高层少 cache(锥形) | arXiv:2406.02069 |
| LagKVPress | Scorer | KV 滞后相关信息,query/attention free | arXiv:2504.04704 |
| KeyDiffPress | Scorer | 基于 key 相似度逐出 token | arXiv:2504.15364 |
| NonCausalAttnPress | Scorer | 非因果分块注意力打分 | arXiv:2507.08143 |
| LeverageScorePress | Scorer | 杠杆分数采样 | arXiv:2507.0???? |
⚠️ LeverageScorePress 论文编号在 README 中显示不完整(
2507.0),建议以 GitHub 源码或 arXiv 搜索确认为准。
§5 Benchmark 数据
官方 HuggingFace Leaderboard 持续更新,以下为关键摘录(Llama 3.x 系列为主):
5.1 精度 vs 压缩比
在 AIME 2024/2025 数学竞赛测试集(KV budget = 2,048 tokens,MATH500 = 512 tokens):
| 方法 | AIME 2024 | AIME 2025 | MATH500 |
|---|---|---|---|
| Full Attention(无压缩) | 57.1% | 40.8% | 69.6% |
| SnapKV | 34.6% | 20.0% | 49.2% |
| R-KV | 25.4% | 17.5% | 46.4% |
| TriAttention | 42.1% | 32.9% | 56.0% |
| ExpectedAttentionPress | (领先) | (领先) | (领先) |
官方图表显示 ExpectedAttentionPress 在各个压缩比下精度均为最高,但具体数值需以 leaderboard 最新结果为准。
5.2 显存节省
Llama 3.1-8B,128K 上下文,50% 压缩:峰值显存从 45 GB 降至 37 GB(-18%),且 decode 速度同步提升。
§6 典型适用场景
- 长文档问答:128K~1M token 文档的 RAG 场景,prefill 显存压力最大的环节。
- Agent 记忆压缩:Agent 系统需要把大量历史上下文压缩后继续推理。
- 长思维链推理:DeepSeek 式千步推理 trace,KV cache 远超单卡容量。
- 多文档摘要:一次性处理大量 PDF / 文本的批量推理。
- 研究者对比评测:快速对比 14+ 种压缩算法在自有数据集上的效果。
§7 坑与注意
⚠️ 坑 1:compression_ratio 针对 context 不针对 question
Pipeline 方式中 press 只压缩 context tokens,question 不参与压缩。如果你的 context 很短(< 1K token),压缩带来的收益几乎为零——此时瓶颈在模型权重而非 KV cache。
⚠️ 坑 2:DecodingPress 是实验性功能
DecodingPress 目前(2025-10)标注为 experimental,压缩间隔(compression_interval)和目标 cache 大小(target_size)需要调参。不兼容所有 press 类型,可能影响输出质量稳定性,生产使用前务必充分测试。
⚠️ 坑 3:不同 press 的压缩质量差异极大
在数学推理等精确任务中,SnapKV 在 12.5% 压缩比下精度从 57.1% 跌到 34.6%(-22.5pp),而 ExpectedAttentionPress 保持显著领先。选错 press + 压缩比可能让模型彻底"失忆"。建议先用 leaderboard 选press,再在自己的数据集上扫压缩比。
⚠️ 坑 4:flash-attention 版本依赖
部分 press(尤其是 ObservedAttentionPress、NonCausalAttnPress)需要 FlashAttention-2 支持,否则报错或退化为普通 attention。安装时加 --extra flash-attn,确认 import flash_attn 不报错。
⚠️ 坑 5:ExpectedAttentionPress 依赖生成分布假设
ExpectedAttentionPress 假设 generation 阶段 query 分布与 prefilling 阶段相似,以此估算未来注意力。若模型在特定任务(如极长思维链)上的 query 分布与训练差异大,可能出现压缩偏误。
⚠️ 坑 6:PyramidKV 与其他 press 不兼容
PyramidKVPress 按层级分配不同压缩比(底层多 cache、高层少),与其他基于全局打分排序的 ScorerPress 逻辑不同,混用会导致行为不符合预期。
§8 与同类对比
| 方案 | 类型 | 优点 | 缺点 |
|---|---|---|---|
| KVPress | 剪枝(pruning) | 14+ 算法统一评测、即装即用、无需训练、支持 transformers pipeline | 精度与压缩比强相关,需选对 press |
| HQQ(HQQLinear) | 量化(quantization) | 模型权重 + KV cache 一起压,精度损失小(< 1%) | 需逐层配置,高压缩比时仍会损失精度 |
| StreamingLLM | 固定 budget | 极简逻辑、确定性的"首尾 token"策略 | 无法感知重要性,高压缩比下精度差 |
| AutoAWQ / GPTQ | 权重量化 | 工业成熟、生态丰富 | 不专门压缩 KV cache,KV 瓶颈仍在 |
| vLLM / SGLang PagedAttention | KV 管理 | 显存分配更高效、batch 利用率高 | 不压缩内容,只是避免碎片化 |
一句话:HQQ/权重量化解决的是"每个值占多少 bit"的问题;KVPress 解决的是"哪些值值得保留"的问题。两者正交,可以叠加使用。
§9 一句话推荐结论
KVPress 是目前 LLM KV cache 压缩领域覆盖算法最全、最贴近工程落地的工具包——无论你是想在大模型上跑长上下文但显存不够,还是研究者需要快速对比 14 种压缩算法,KVPress 都值得优先上手。其中 ExpectedAttentionPress 在精度/压缩比上整体领先,首次使用建议从它和
compression_ratio=0.5开始。
数据来源:GitHub README(NVIDIA/kvpress,2025-10-01 fetch)、HuggingFace Blog: Mastering Long Contexts in LLMs with KVPress、NVIDIA Research Blog: KV Cache Compression and Its Infra Problems、arXiv:2510.00636v1(Expected Attention)。DecodingPress 为实验性功能,相关参数以 GitHub 最新版为准。