Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs
- 类型:arxiv
- 标识:2608.20953
- 链接:https://arxiv.org/abs/2608.20953
- 主分类:llm-infra
- 形态:application
- TLDR:Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding, and long-context behavior enough to require a recovery, or healing, stage before deployment. The default recipe, quantization-aware training (QAT), re-fits the compressed, quantized model to hard labels; in our pipeline it converged slowly and collapsed past its peak. We adopted Quantization-Aware Healing (QAH) instead. Because a structurally compressed model is neve
- 副分类:engineering
- 待LLM分类:否
- 标题中文:量化感知修复:恢复压缩后 4-Bit LLM 的实用方案
- TLDR中文:低成本服务大语言模型日益意味着交付既经过结构压缩至部分参数、又量化到 4 bit 的模型。这两个步骤叠加后会显著降低推理、数学、代码以及长上下文能力,因此在部署前需要一个恢复(即修复)阶段。默认方案 quantization-aware training (QAT) 将压缩并量化后的模型重新拟合到硬标签;在我们的流水线中它收敛缓慢,且在峰值之后出现性能崩塌。我们转而采用 Quantization-Aware Healing (QAH)。由于结构压缩后的模型从未
- 来源文件:
- /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json