评估 vLLM 在分离式 serving 下的精度,并通过训练后量化在最高 2.8T 参数的模型上进一步验证共享权重格式的分离式 serving。Assessment of accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters are evaluated.
论文
151 张论文卡片 · LLM 基础设施 · 方法
NetKV,一种使用该 oracle 信息的 O(|D|) 每请求贪心策略,其层级排序被证明对过时遥测数据具有鲁棒性;并证明随着上下文长度增长,忽略网络项会使仅缓存感知的调度任意次优。NetKV, the O(|D|) per-request greedy that consumes this oracle, has tier rankings that are provably robust to stale telemetry, and it is proved that ignoring the network term renders cache-aware-only scheduling arbitrarily suboptimal as context length grows.
受 Apache Spark 等分布式计算框架启发,DISCO 将长上下文切分到一组专职 Worker LLM 上,进行并行、局部的 grounding,建立了鲁棒长上下文推理的高效范式。Inspired by distributed computing frameworks like Apache Spark, DISCO partitions long context across a fleet of Worker LLMs dedicated exclusively to parallel, localized grounding, establishing a highly efficient paradigm for robust long-context inference.
本工作提出 LMCACHE,首个也是目前最高效的开源 KV 缓存方案,可将现代 LLM 引擎生成的 KV 缓存从 GPU 显存中提取并存储,并跨引擎和查询共享。This work presents LMCACHE, the first and so far the most efficient open-source KV caching solution, which extracts and stores KV caches generated by modern LLM engines out of the GPU memory and shares them across engines and queries.
提出 G²PTQ,一个具有广义梯度补偿的统一 PTQ 框架,在全局监督的分块优化目标下融合一阶与二阶信息,能更好地对齐全精度模型,性能优于 SOTA 基线。G$^2$PTQ is presented, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective and enables better alignment with the full-precision model, outperforming state-of-the-art baselines.
通道级 history-state 注入结合 timestep 条件化构成了一种低成本且更有效的设计,能够在更长 unrolling 下更好地保留知识,同时提升在不同推理预算下的鲁棒性。It is shown that channel-wise history-state injection combined with timestep conditioning offers a low-cost and more effective design, better preserving knowledge under extended unrolling while improving robustness across inference budgets.
本文提出 FRAC,一种源自分数阶动力学的选择性 SSM 架构,用幂律长记忆替代指数衰减,在长上下文性能上持续优于 SOTA SSM 基线,同时在短上下文上保持竞争力。FRAC is introduced, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory and consistently improves long-context performance over state-of-the-art SSM baselines while staying competitive on short-context.
本文提出 Loop Scaling Laws,是首个将 recurrence 和 sparsity 与模型规模、数据联合建模的 scaling law,为在算力和显存约束下设计 looped MoE 模型提供了原则性基础。Loop Scaling Laws are introduced, the first scaling law to jointly model recurrence and sparsity alongside model size and data, and provide a principled foundation for designing looped MoE models under compute and memory constraints.
基于 RoPE 的语言模型出现长上下文失效,根源在于 RoPE 在维持稳定 token 偏好与区分相近位置之间存在固有权衡。要判断应处理哪种弱点以及如何处理,需要更精确地刻画 RoPE 在不同上下文长度下训练后模型中的行为。我们通过允许 RoPE 各频率下 query-key 尺度不一致,弥补了先前理论的一个关键局限,使之与实际经验观测高度吻合。我们的理论使得上述脆弱性对单个注意力头与输入可测量,并量化了高频Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns well with practical empirical observations. Our theory makes both vulnerabilities measurable for individual heads and inputs, and quantifies how high-freque
结果表明,有效的拒绝能够保留任务结构,同时限制与参考策略的耦合,并且较小的冻结模型可以低成本地提供此类拒绝能力。The results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.
本文提出 MemFold,根据所支持的行为来优化固定预算的软记忆;在 PersonaMem-32K 和 PersonaMem-128K 上取得了作者所测得的最高准确率,且在更长历史长度下优势进一步扩大。This work presents MemFold, which optimizes a fixed-budget soft memory by the behavior it supports, and attains the highest accuracy the authors measure on PersonaMem-32K and PersonaMem-128K, with margins that widen at the longer history length.
提出 FOVEATED,一种即插即用框架,通过随机偏移分配给前文上下文 key 的 Rotary Position Embedding 位置来构建每个句子的聚焦视图,并从理论上分析了 FOVEATED 如何抵消上下文引发的难度低估。FOVEATED, a plug-and-play framework that constructs focused views of each sentence by randomly shifting the Rotary Position Embedding positions assigned to the keys of its preceding context, is proposed and theoretically analyzed how FOVEATED counteracts context-induced difficulty underestimation is analyzed.
Extender 在 199M–924M 参数规模下匹配 Transformer 在短上下文(CORE)任务上的准确率,并在 924M 参数下于长上下文(RULER)任务上超越 Transformer。The Extender matches Transformer accuracy on short-context (CORE) tasks at 199M-924M parameters, and exceeds Transformer accuracy on long-context (RULER) workloads, again at 924M parameters.
提出 Behavior-Preserving KV Cache Compression,一种无需训练的框架,通过估计候选条目被移除后压缩缓存所诱导的 logits,并利用驱逐前的前向统计量评估其与完整缓存下一 token 分布之间的 KL 散度来为候选驱逐打分。This work proposes Behavior-Preserving KV Cache Compression, a training-free framework that scores candidate evictions by estimating the compressed-cache logits induced by their removal and evaluating the resulting KL to the full-cache next-token distribution using pre-eviction forward statistics.
结果表明,基于查询的循环记忆组合能够提升长上下文建模能力,并在超出训练上下文范围之后依然有效,同时每个 chunk 仅使用紧凑的仿射摘要。Results demonstrate that query-dependent composition of recurrent memory improves long-context modeling and remains effective beyond the training context while using compact per-chunk affine summaries.
提出基于流的自潜在推理(Flow-based Latent Reasoning, FLaRe),一份简洁方案,涵盖潜空间编码内容及其塑形方式、流训练位置、答案读取方式,以及最终在模型自身已验证思考上进行训练的阶段。Flow-based Latent Reasoning (FLaRe), a simple recipe covering what the latent space encodes and how to shape it, where to train the flow, how to read out the answer, and a final stage of training on the model's own verified thoughts, is presented.
大语言模型只能使用其上下文窗口中容纳的文本,并且每次发送 prompt 时都会重新计算其内部的 key-value (KV) 状态。我们测试了一个记忆层,即公开包 galahad-kv,它将每个约 16000 token 块的 KV 状态保存到加密的本地 NVMe 磁盘,并在之后按字节精确、无重计算地加载回来。我们在 5000 万 token 的真实公共文本上运行,通过单块 NVIDIA H100 上的 vLLM 提供服务,配合 Gemma 4 12B 和 Gemma 4 31B。我们探测的每个块都从加密存储中无重计算地加载回来(100 个中的 100 个A large language model can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent. We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it. We ran it on 50,000,000 tokens of real public text, served through vLLM on one NVIDIA H100, with Gemma 4 12B and Gemma 4 31B. Every block we probed was loaded back from the encrypted store with no recompute (100 of 100,
这是一本关于大语言模型的书。正如标题所示,本书主要聚焦于基础概念,而非全面覆盖所有前沿技术。全书共分六个主要章节,每章探讨一个关键领域:预训练、生成模型、提示工程、对齐、推理和推理能力。本书面向自然语言处理及相关领域的大学生、从业者和研究者,也可作为任何对大语言模型感兴趣的人的参考。This is a book about large language models. As indicated by the title, it primarily focuses on foundational concepts rather than comprehensive coverage of all cutting-edge technologies. The book is structured into six main chapters, each exploring a key area: pre-training, generative models, prompting, alignment, inference, and reasoning. It is intended for college students, professionals, and practitioners in natural language processing and related fields, and can serve as a reference for anyone interested in large language models.
提出 KVpop,通过直接监督保留或丢弃决策来学习固定预算的 KV 淘汰策略,并引入基于延迟记忆的打分器,在已学习的淘汰方法中独有地延迟固定步数的打分,以利用近期上下文。KVpop is introduced, which learns a fixed-budget KV eviction policy by directly supervising the keep-or-drop decision, and introduces a delayed memory-based scorer that, uniquely among learned eviction methods, defers scoring for a fixed number of steps to exploit near-future context.
提出流形约束假设(Manifold Constraint Hypothesis, MCH):若自然表征集中于一个结构化的低维流形上,则干预应被约束在该流形内,从而在干预过程中更好地保留表征中编码的其他信息。The Manifold Constraint Hypothesis (MCH) is proposed: if natural representations concentrate on a structured, lower-dimensional manifold, then interventions should be constrained to that manifold and better preserve other information encoded in the representation during interventions.
提出 Simplified Sparse Attention——一种更简洁的稀疏注意力方案,无需任何架构改动,在检索增强生成中优于全注意力;并扩展为分层 gist-of-gist 变体,在高达 32× 的高压缩比下保持或提升精度,同时实现对数级解码复杂度。Simplified Sparse Attention is introduced, a simpler approach to sparse attention that requires no architectural changes that outperforms full attention in retrieval-augmented generation and extends to a hierarchical gist-of-gist variant that achieves log-linear decoding complexity while maintaining or improving accuracy at high compression ratios up to 32x.
提出 Erase-then-Delta Attention (EDA),一种将"在哪里擦除"与"在哪里写入"解耦的内存更新规则;研究表明循环记忆模型不仅应决定写入什么,还应决定擦除哪些陈旧信息以及擦除的位置。Erase-then-Delta Attention (EDA), a memory update rule that decouples where to erase from where to write, is proposed, suggesting that recurrent memory models should decide not only what to write, but also what stale information to erase and where.
本文从前瞻视角重新审视 token 重要性,提出衡量压缩 token 对未来上下文影响的新指标 Forward Influence,以及融合信息论信号的熵感知 KV cache 压缩框架 InfoKV。This paper revisits token importance from a forward-looking perspective and introduces Forward Influence, a metric that measures how compressed tokens affect future contexts, and proposes InfoKV, an entropy-aware KV cache compression framework that incorporates information-theoretic signals.
提出 SAC,这是首个针对稀疏注意力模型优化的高效解耦 KV Cache 系统,利用 CXL(Compute Express Link)的低延迟、cache-line 粒度 load/store 语义,将基于 CXL 的解耦确立为新兴稀疏注意力模型的优越基础设施。SAC is proposed, the first efficient disaggregated KV cache system optimized for sparse attention models, which leverages the low-latency, cache-line granularity load/store semantics of Compute Express Link (CXL), establishing CXL-based disaggregation as the superior infrastructure for emerging sparse attention models.
提出 Sparse Delta Memory,一种通过稀疏寻址方案将门控线性 RNN 隐状态容量扩展数个数量级的架构,在上下文学习与长上下文检索任务上显著提升性能。Sparse Delta Memory is introduced, an architecture that scales the hidden state of gated linear RNNs to orders of magnitude higher capacity using a sparse addressing scheme and significantly improves performance on in-context learning and long-context retrieval tasks.
本文对 softmax 注意力与四种近期的循环线性注意力架构(DeltaNet、Gated DeltaNet、Kimi Delta Attention 与 Gated DeltaNet-2)进行对比研究,明确阐述它们在表达能力、记忆衰减、擦写控制、训练吞吐量与实现复杂度上的差异。A comparative study of softmax attention and four recent recurrent linear-attention architectures: DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2 is presented, making explicit how they differ in expressivity, memory decay, erase and write control, training throughput, and implementation complexity.
研究表明模型具备相关子群体知识,但难以稳定传递到聚合估计中,这一差距使统计自一致性成为评估 LLM 的尚未饱和、无需参考的准则。It is suggested that models possess relevant subpopulation knowledge but do not reliably propagate it into aggregate estimates, and this gap establishes statistical self-consistency as an unsaturated, reference-free criterion for evaluating LLMs.
提出 BloombergGPT,一个 500 亿参数的语言模型,在广泛的金融数据上训练而成,并基于 Bloomberg 丰富的数据源构建了包含 3630 亿 token 的数据集,可能是迄今最大的领域专用数据集。This work presents BloombergGPT, a 50 billion parameter language model that is trained on a wide range of financial data, and constructs a 363 billion token dataset based on Bloomberg's extensive data sources, perhaps the largest domain-specific dataset yet.
综合评估表明,DeepSeek-V3 优于其他开源模型,并达到与领先闭源模型相当的性能。Comprehensive evaluations reveal that DeepSeek-V3 outperforms other open-source models and achieves performance comparable to leading closed-source models.
这些发现表明,即便长上下文评估正从简单检索转向复杂推理,对相关证据的准确 grounding 仍是一项不可或缺且仍有大幅提升空间的能力。These findings suggest that, even as long-context evaluation shifts from simple retrieval toward complex reasoning, accurate grounding in relevant evidence remains an indispensable capability with substantial room for improvement.
本文描述了 TensorFlow 接口及 Google 构建的该接口实现,已被用于开展研究,并在计算机科学及其他十余个领域中将机器学习系统部署至生产环境。The TensorFlow interface and an implementation of that interface that is built at Google are described, which has been used for conducting research and for deploying machine learning systems into production across more than a dozen areas of computer science and other fields.
本文推出参数规模从 7B 到 65B 的基础语言模型集合 LLaMA,并证明完全使用公开数据集即可训练出 SOTA 模型,无需依赖专有或不可获取的数据。LLaMA, a collection of foundation language models ranging from 7B to 65B parameters, is introduced and it is shown that it is possible to train state-of-the-art models using publicly available datasets exclusively, without resorting to proprietary and inaccessible datasets.
通过从概率分布的动态 nucleus 中采样文本,可在有效截断不可靠分布尾部的同时保持多样性,使生成文本更接近人类文本质量,在不牺牲流畅性与连贯性的前提下提升多样性。By sampling text from the dynamic nucleus of the probability distribution, which allows for diversity while effectively truncating the less reliable tail of the distribution, the resulting text better demonstrates the quality of human text, yielding enhanced diversity without sacrificing fluency and coherence.
本文指出基于 Transformer 的次二次时间模型的关键缺陷在于无法执行基于内容的推理,并将选择性 SSM 集成到不包含注意力乃至 MLP 块的简化端到端神经网络架构(Mamba)中。This work identifies that a key weakness of subquadratic-time models based on Transformer architecture is their inability to perform content-based reasoning, and integrates selective SSMs into a simplified end-to-end neural network architecture without attention or even MLP blocks (Mamba).
该设计将 hypernetwork 的注入能力与目标模型的通用能力解耦,首次实现了对 hypernetwork 架构 scaling law 的严格研究,并提供了首个基于实证的 scaling law,用以指导大语言模型中面向事实推理的 hypernetwork 设计。The design decouples the hypernetwork's injection capacity from the target model's general capability, enabling, for the first time, a rigorous study of scaling laws for hypernetwork architectures, and provides the first empirically grounded scaling laws to guide hypernetworks for factual reasoning in large language models.