Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
- 类型:arxiv
- 标识:2608.09444
- 链接:https://arxiv.org/abs/2608.09444
- 主分类:llm-infra
- 形态:method
- TLDR:A main promise of looped language models is depth-adaptive inference. By looping a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. However, tokens with different numbers of loops cannot share a uniform forward pass and therefore cannot be handled by standard batching systems such as vLLM. The practical value of depth-adaptive inference thus hinges on whether batching can be made efficient. We introduce the first efficient method for depth-adaptive looped LMs via continuous depth batching (CDB), which forms new batche
- 待LLM分类:否
- 标题中文:[标题中文] 通过连续深度批处理实现循环语言模型的深度自适应推理
- TLDR中文:[TLDR中文] 循环语言模型的一个主要承诺是深度自适应推理。通过将一组共享层循环可变次数,模型可对"简单" token 使用更少算力,对"困难" token 使用更多算力。然而,具有不同循环次数的 token 无法共享统一的 forward pass,因此无法由 vLLM 等标准批处理系统处理。因此,深度自适应推理的实际价值取决于批处理能否变得高效。我们提出首个通过连续深度批处理(CDB)实现循环语言模型深度自适应的有效方法,该方法在新批次的组
- 来源文件:
- /inbox/tom/_candidates/2026-09-28-agent-rag-longcontext-candidates.json