Verification-Aware Training for Speculative Decoding

  • 类型:arxiv
  • 标识:2608.30135
  • 链接:https://arxiv.org/abs/2608.30135
  • 主分类:engineering
  • 形态:position
  • TLDR:Speculative decoding accelerates large language model inference by using a draft model to generate candidate tokens, which are verified by the target model in a single forward pass. Verification proceeds sequentially and discards every position from the first rejection onward, yet existing draft training relies on token-level imitation of the target with a fixed per-position weighting that reflects neither property. We introduce Verification-Aware Training (VAT), a plug-in framework that simulates verification at every training step and turns the resulting accept and reject patterns into super
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-01-agent-rag-longcontext-candidates.json