G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
- 类型:arxiv
- 标识:2609.31009
- 链接:https://arxiv.org/abs/2609.31009
- 主分类:llm-infra
- 形态:method
- TLDR:Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise objectives lack global supervision; while methods with global objectives fix their Hessian estimates at the start and ignore first-order gradients, so their guidance grows stale as quantization proceeds. This paper presents G^2PTQ, a unified PTQ framework with Generalized Gradient Compensation that int
- 副分类:engineering
- 待LLM分类:否
- 标题中文:G^2PTQ:通过广义梯度补偿改进 LLM 训练后量化
- TLDR中文:训练后量化(PTQ)是一种无需重新训练即可降低大语言模型(LLM)内存与计算开销的实用方法。基于 GPTQ 的方法已成为事实上的标准,但它们面临两个互补的局限:采用局部逐层目标的方法缺乏全局监督;而采用全局目标的方法则在开始时就固定了 Hessian 估计,并忽略了一阶梯度,导致其指导信号随着量化推进而逐渐失效。本文提出 G^2PTQ——一个通过广义梯度补偿将两者统一的 PTQ 框架。
- 来源文件:
- /inbox/tom/_candidates/2026-09-29-agent-rag-longcontext-candidates.json