How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus

  • 类型:arxiv
  • 标识:2609.15504
  • 链接:https://arxiv.org/abs/2609.15504
  • 主分类:llm-infra
  • 形态:method
  • TLDR:Orthrus is a hybrid autoregressive-diffusion architecture that accelerates autoregressive language-model inference by generating multiple tokens in parallel while using a frozen autoregressive backbone. Its central claim is that an intra-model consensus mechanism enables lossless speculative decoding, producing the same output sequence as the autoregressive model. We independently reproduce Orthrus and examine this claim under different numerical precisions. Under BF16 inference, exact trajectory matching occurs in only 45% of cases for the authors' checkpoint and 43% for our independently tra
  • 待LLM分类:否
  • 标题中文:无损推测解码真的无损吗?数值精度在 Orthrus 中的作用
  • TLDR中文:Orthrus 是一种混合自回归-扩散架构,通过并行生成多个 token 并使用冻结的自回归骨干来加速自回归语言模型推理。其核心声明是模型内共识机制可实现无损推测解码,输出与自回归模型相同的序列。我们独立复现了 Orthrus,并在不同数值精度下检验该声明。在 BF16 推理下,作者的 checkpoint 仅在45%的情况下精确匹配轨迹,我们独立训练的 checkpoint 则为43%
  • 来源文件
  • /inbox/tom/_candidates/2026-09-15-agent-rag-longcontext-candidates.json