Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

  • 类型:arxiv
  • 标识:2608.20953
  • 链接:https://arxiv.org/abs/2608.20953
  • 主分类:llm-infra
  • 形态:application
  • TLDR:Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding, and long-context behavior enough to require a recovery, or healing, stage before deployment. The default recipe, quantization-aware training (QAT), re-fits the compressed, quantized model to hard labels; in our pipeline it converged slowly and collapsed past its peak. We adopted Quantization-Aware Healing (QAH) instead. Because a structurally compressed model is neve
  • 副分类:engineering
  • 待LLM分类:否
  • 标题中文:量化感知修复:恢复压缩后 4-Bit LLM 的实用方案
  • TLDR中文:低成本服务大语言模型日益意味着交付既经过结构压缩至部分参数、又量化到 4 bit 的模型。这两个步骤叠加后会显著降低推理、数学、代码以及长上下文能力,因此在部署前需要一个恢复(即修复)阶段。默认方案 quantization-aware training (QAT) 将压缩并量化后的模型重新拟合到硬标签;在我们的流水线中它收敛缓慢,且在峰值之后出现性能崩塌。我们转而采用 Quantization-Aware Healing (QAH)。由于结构压缩后的模型从未
  • 来源文件
  • /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json