G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation

  • 类型:arxiv
  • 标识:2609.31009
  • 链接:https://arxiv.org/abs/2609.31009
  • 主分类:llm-infra
  • 形态:method
  • TLDR:Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise objectives lack global supervision; while methods with global objectives fix their Hessian estimates at the start and ignore first-order gradients, so their guidance grows stale as quantization proceeds. This paper presents G^2PTQ, a unified PTQ framework with Generalized Gradient Compensation that int
  • 副分类:engineering
  • 待LLM分类:否
  • 标题中文:G^2PTQ:通过广义梯度补偿改进 LLM 训练后量化
  • TLDR中文:训练后量化(PTQ)是一种无需重新训练即可降低大语言模型(LLM)内存与计算开销的实用方法。基于 GPTQ 的方法已成为事实上的标准,但它们面临两个互补的局限:采用局部逐层目标的方法缺乏全局监督;而采用全局目标的方法则在开始时就固定了 Hessian 估计,并忽略了一阶梯度,导致其指导信号随着量化推进而逐渐失效。本文提出 G^2PTQ——一个通过广义梯度补偿将两者统一的 PTQ 框架。
  • 来源文件:
  • /inbox/tom/_candidates/2026-09-29-agent-rag-longcontext-candidates.json