Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching

  • 类型:arxiv
  • 标识:2608.09444
  • 链接:https://arxiv.org/abs/2608.09444
  • 主分类:llm-infra
  • 形态:method
  • TLDR:A main promise of looped language models is depth-adaptive inference. By looping a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. However, tokens with different numbers of loops cannot share a uniform forward pass and therefore cannot be handled by standard batching systems such as vLLM. The practical value of depth-adaptive inference thus hinges on whether batching can be made efficient. We introduce the first efficient method for depth-adaptive looped LMs via continuous depth batching (CDB), which forms new batche
  • 待LLM分类:否
  • 标题中文:[标题中文] 通过连续深度批处理实现循环语言模型的深度自适应推理
  • TLDR中文:[TLDR中文] 循环语言模型的一个主要承诺是深度自适应推理。通过将一组共享层循环可变次数,模型可对"简单" token 使用更少算力,对"困难" token 使用更多算力。然而,具有不同循环次数的 token 无法共享统一的 forward pass,因此无法由 vLLM 等标准批处理系统处理。因此,深度自适应推理的实际价值取决于批处理能否变得高效。我们提出首个通过连续深度批处理(CDB)实现循环语言模型深度自适应的有效方法,该方法在新批次的组
  • 来源文件:
  • /inbox/tom/_candidates/2026-09-28-agent-rag-longcontext-candidates.json