Unlocking Lossless Speedups in LLMs via Discrete Diffusion
- 类型:arxiv
- 标识:2609.04010
- 链接:http://arxiv.org/abs/2609.04010v1
- 主分类:multimodal
- 形态:method
- TLDR:Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution. We decouple the parameters of these models into two sets: AR weights, trained using the standard NTP objective, and lightweight diffusion weights, trained to generate multiple tokens simultaneously. The diffusion weight
- 待LLM分类:否
- 标题中文:通过离散扩散释放 LLM 的无损加速
- TLDR中文:大语言模型(LLM)的成功很大程度上归功于 next-token prediction(NTP),但其自回归(AR)结构需要缓慢的串行 token 生成。为克服这一瓶颈,我们提出扩散增强 LLM,一类新模型,在使用扩散从该分布中并行采样多个 token 的同时定义 AR 模型分布。我们将这些模型的参数解耦为两组:AR 权重,使用标准 NTP 目标训练;轻量扩散权重,训练用于同时生成多个 token。扩散权重
- 来源文件:
- /inbox/tom/_candidates/2026-09-08-agent-rag-longcontext-candidates.json