Pretraining Transformers with Quantized Softmax in Attention

  • 类型:arxiv
  • 标识:2609.33591
  • 链接:https://arxiv.org/abs/2609.33591
  • 主分类:engineering
  • 形态:method
  • TLDR:Low-precision Transformer systems increasingly quantize attention matrix multiplications, while softmax often remains at higher precision. During pretraining, an approximate softmax changes the gradients that train the model as well as its forward computation. We study this interaction with K-interval attention, which approximates the exponential using K+1 grid values. We vary per-row grid calibration, interpolation versus hard rounding, and the placement of a straight-through surrogate relative to normalization. We derive the corresponding backward rules, including calibration derivatives, an
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-09-30-agent-rag-longcontext-candidates.json