Disaggregated Quantization: Specializing LLM Prefill and Decode

  • 类型:arxiv
  • 标识:2609.26333
  • 链接:https://arxiv.org/abs/2609.26333
  • 主分类:llm-infra
  • 形态:method
  • TLDR:Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to both of these phases. On Qwen 3 and Gemma 3, removing activation quantization specifically on decode improves accuracy on decode-heavy tasks without increasing inference cost. Training separate compute-native prefill weights accelerates prompt processing relative to weight-only inference while matc
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-09-29-rag-retrieval-reranking-candidates.json
  • /inbox/tom/_candidates/2026-09-29-agent-rag-longcontext-candidates.json