IntBMoE 是一种 block-conditioned MoE,将三者解耦,通过密集专家组合与稀疏 block 执行配对实现;在图像分类任务上,相较代表性的稀疏与密集 MoE 基线均取得稳定提升。IntBMoE is a block-conditioned MoE that decouples all three by pairing dense expert composition with sparse block execution, and shows consistent gains over representative sparse and dense MoE baselines on image classification.
论文
187 张论文卡片 · LLM 基础设施
本文提出 TTKV,一种 KV cache 管理框架,将人类记忆系统映射到具备异构容量与精度的 KV cache 上,在 128K 上下文任务上将跨层流量降低 5.94 倍。TTKV is proposed, a KV cache management framework that maps the human memory system onto the KV cache with heterogeneous capacity and precision, and reduces cross-tier traffic by 5.94x on 128K-context tasks.
在更大规模 dense 模型与 MoE 模型上的比较进一步凸显了在持续生成过程中调度 prompt 处理的重要性:在并发负载下,vllm-metal 的 packed prefill-decode 路径维持了低于 omlx 的首 token 延迟。Comparisons on larger dense and MoE models reinforce the importance of scheduling prompt processing alongside ongoing generation: vllm-metal's packed prefill-decode path maintains lower first-token latency than omlx under concurrent load.
表明 Kimi Delta Attention (KDA) 可通过单次 delta-rule 变换与其通道门控提供的反射组合实现 2D 旋转,并刻画 CKDA 的表达能力,证明每个正交的对角加秩-1矩阵恰为一个 CKDA 转移矩阵。This work shows that Kimi Delta Attention (KDA) can realize 2D rotations by combining a single delta-rule transformation with a second reflection supplied by its channel-wise gate, and characterize the expressivity of CKDA and prove that every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition matrix.
本文提出 KV Policy (KVP),一种仅基于 key 与 value 向量、在预计算生成轨迹上训练的轻量级 per-head RL Agent 框架,证明学习预测未来 token 效用是自适应 KV cache 管理中强大且可扩展的范式。KV Policy (KVP), a framework of lightweight per-head RL agents trained on pre-computed generation traces using only key and value vectors, is introduced, demonstrating that learning to predict future token utility is a powerful and scalable paradigm for adaptive KV cache management.
政治话语日益呈现更强的敌对感是普遍观感,但稳健证据仍稀缺。本文研究 1997 至 2026 年丹麦议会的归责行为,结合专门构建的分类器 BlameBERT(F1: 0.80)与多层统计建模。该分类器采用面向低至中等资源语言的标注高效流程。结果显示出一条香蕉形轨迹:归责水平约在 2016 年前下降,随后在近年(2019–2026)进入显著且持续的上升阶段。执政地位显著Political discourse is widely perceived to be growing more hostile, yet robust evidence remains scarce. This study examines blame attribution in the Danish Parliament from 1997 to 2026, combining a purpose-built classifier, BlameBERT (F1: 0.80), with multilevel statistical modeling. The classifier is constructed using an annotation-efficient pipeline for blame attribution in low-to-mid resource languages. The results reveal a banana-shaped trajectory, with blame declining until around 2016 before entering a significant and sustained increase in recent years (2019-2026). Government status consi
本文展示了在一对架构匹配的 Qwen3.5 4B-to-9B 兄弟模型之间实现有用的持久 hybrid-state transfer,是首次在大小不同的混合语言模型之间完成持久循环推理状态的跨模型交接,且无需 target prefix replay。This work demonstrates useful persistent hybrid-state transfer across one architecture-matched Qwen3.5 4B-to-9B sibling pair, the first demonstrated cross-model handoff of persistent recurrent inference state between differently sized hybrid language models without target prefix replay.
TritonForge 是一个面向自动化 Triton kernel 优化的 profiling 引导框架,融合 kernel 分析、运行时 profiling 与迭代式代码转换以简化优化流程,并为自动化 GPU 性能优化领域的未来研究奠定基础。TritonForge, a profiling-guided framework for automated Triton kernel optimization that integrates kernel analysis, runtime profiling, and iterative code transformation to streamline the optimization process and provides a foundation for future research in automated GPU performance optimization.
本文探讨容器化、微服务、动态调度等 Cloud Native 技术如何从根本上提升 LLM 推理服务,并展示 Cloud Native 系统在高需求场景下实现更高效资源分配、降低延迟与提升吞吐的能力。This article explores how Cloud Native technologies, such as containerization, microservices, and dynamic scheduling, can fundamentally improve LLM inference serving and demonstrates how a Cloud Native system enables more efficient resource allocation, reduces latency, and enhances throughput in high-demand scenarios.
本文提出一种方法,通过逐层剔除后 WER(Word Error Rate)的变化对编码器层进行排序,并给出剪枝后的模型——即一个层数更少的更浅编码器。This work presents an approach that ranks encoder layers by the leave-one-layer-out change in Word Error Rate (WER), and presents the pruned model, which is simply a more shallow encoder with fewer layers.
本文论证线性表征假设不是单一假设,而是一组通过表征等价性加以区分的命题族;该框架厘清了假设如何随度量、读取点与分析阶段变化,并被用于审计常见表征量与近期可解释性分析。It is argued that the Linear Representation Hypothesis is not one hypothesis but a family of claims distinguished by representation equivalence, and this framework clarifies how assumptions can change across metrics, reading points, and analysis stages, and is used to audit common representation quantities and recent interpretability analyses.
本文首次对 vLLM 启动延迟进行了详细的性能表征,并构建了一个轻量级分析模型,能够针对给定硬件配置准确预测 vLLM 的启动延迟,为大规模推理环境中的资源规划提供了可操作的指导。This paper presents the first detailed performance characterization of vLLM startup latency and develops a lightweight analytical model that accurately predicts vLLM's startup latency for a given hardware configuration, providing actionable guidance for resource planning in large-scale inference environments.
本文提出 Neural Spectral Capacity (NSC),一种以每个权重矩阵的奇异值谱为依据的闭式标量,在七类 Transformer 与 CNN 系列上的排序效果优于 #Params、#FLOPs 及代表性免训练代理指标。This work proposes Neural Spectral Capacity (NSC), a closed-form scalar grounded in the singular-value spectrum of each weight matrix, which outperforms #Params, #FLOPs, and representative training-free proxies in ranking across seven Transformer and CNN families.
本文提出 PISA,一种采用金字塔式 Top-K 选择策略的 block-sparse attention 机制,并为训练与推理开发了硬件感知的 Triton kernel,将层级路由与 LogSumExp 评分融合,无需显式构造 query-key score 矩阵。PISA is proposed, a block-sparse attention mechanism that employs a pyramid Top-$K selection strategy, and develops hardware-aware Triton kernels for both training and inference, fusing hierarchical routing and LogSumExp scoring without materializing the query-key score matrix.
本文提出 PsPLUG,一个轻量级插件式模块,在刻画目标风格后学习用户专属的残差,能更好地保留用户偏好,并对个性化与风格遵循之间的平衡提供更精确的控制。PsPLUG is proposed, a lightweight plug-in that learns a user-specific residual after accounting for the requested style, and better preserves user preferences while providing precise control over the balance between personalization and style adherence.
本文提出首个通过 continuous depth batching (CDB) 实现 depth-adaptive looped LM 的高效方法:在 loop 步骤之间重组 batch,动态调度架构中的 looped 与非 looped 部分,管理 looped KV-caching,并提前预测将退出 loop 的 token,以便异步准备 batch。This work introduces the first efficient method for depth-adaptive looped LMs via continuous depth batching (CDB), which forms new batches between loop steps, and dynamically schedules looped and non-looped parts of the architecture, manages looped KV-caching, and predicts which tokens will exit the loop in advance so it can prepare batches asynchronously.
使用一种分层旋转位置编码(hRoPE),将段落、句子和 token 索引表示为独立通道,保持 token 序列固定,对段落坐标 $p_1$ 进行干预,并使用 token 距离精确估计器测量跨段落注意力。A hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate $p_1$, and measure cross-paragraph attention with a token-distance-exact estimator is used.
Boxoffice 是一个程序化生成评测数据集的工具,用以覆盖具有挑战性的 KV cache 复用模式;研究表明现有数据集并未展现出充分评估此类技术所需的复用动态特性。Boxoffice, a tool that programmatically generates evaluation datasets that exercise challenging KV cache reuse patterns, is introduced and it is shown that existing datasets do not exhibit the reuse dynamics needed to thoroughly evaluate such techniques.
Intent2Tc 是一个由语言模型驱动的闭环框架,将业务级流量整形意图转换为声明式子意图,进而生成经过验证的可执行 Linux traffic control (tc) 配置,并展示了该框架的实际适用性。Intent2Tc is presented, a closed-loop language-model-driven framework that translates business-level traffic-shaping intents into declarative sub-intents and subsequently into validated, executable Linux traffic control (tc) configurations and demonstrates the practical applicability of the proposed framework.
本文提出 softmax reparameterization,一种训练后方法,在量化前搜索功能等价的输出头,并展示了在总体 logit 误差增大的情况下保真度仍可提升。This work introduces softmax reparameterization, a post-training method that searches over functionally equivalent output heads before quantization and shows how fidelity can improve despite greater total logit error.
评估 vLLM 在分离式 serving 下的精度,并通过训练后量化在最高 2.8T 参数的模型上进一步验证共享权重格式的分离式 serving。Assessment of accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters are evaluated.
NetKV,一种使用该 oracle 信息的 O(|D|) 每请求贪心策略,其层级排序被证明对过时遥测数据具有鲁棒性;并证明随着上下文长度增长,忽略网络项会使仅缓存感知的调度任意次优。NetKV, the O(|D|) per-request greedy that consumes this oracle, has tier rankings that are provably robust to stale telemetry, and it is proved that ignoring the network term renders cache-aware-only scheduling arbitrarily suboptimal as context length grows.
受 Apache Spark 等分布式计算框架启发,DISCO 将长上下文切分到一组专职 Worker LLM 上,进行并行、局部的 grounding,建立了鲁棒长上下文推理的高效范式。Inspired by distributed computing frameworks like Apache Spark, DISCO partitions long context across a fleet of Worker LLMs dedicated exclusively to parallel, localized grounding, establishing a highly efficient paradigm for robust long-context inference.
本工作提出 LMCACHE,首个也是目前最高效的开源 KV 缓存方案,可将现代 LLM 引擎生成的 KV 缓存从 GPU 显存中提取并存储,并跨引擎和查询共享。This work presents LMCACHE, the first and so far the most efficient open-source KV caching solution, which extracts and stores KV caches generated by modern LLM engines out of the GPU memory and shares them across engines and queries.
提出 G²PTQ,一个具有广义梯度补偿的统一 PTQ 框架,在全局监督的分块优化目标下融合一阶与二阶信息,能更好地对齐全精度模型,性能优于 SOTA 基线。G$^2$PTQ is presented, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective and enables better alignment with the full-precision model, outperforming state-of-the-art baselines.
通道级 history-state 注入结合 timestep 条件化构成了一种低成本且更有效的设计,能够在更长 unrolling 下更好地保留知识,同时提升在不同推理预算下的鲁棒性。It is shown that channel-wise history-state injection combined with timestep conditioning offers a low-cost and more effective design, better preserving knowledge under extended unrolling while improving robustness across inference budgets.
本文揭示使用 chunked KV-cache 压缩的模型中存在一种系统性不对称:相同信息在一个阶段易于检索,在另一阶段却难以获取,暴露出平均 benchmark 分数可能掩盖的周期性薄弱点。A systematic asymmetry in models using chunked KV-cache compression is uncovered: the same information can be easy to retrieve at one phase and difficult at another, revealing periodic weak spots that average benchmark scores can conceal.
本文提出 FRAC,一种源自分数阶动力学的选择性 SSM 架构,用幂律长记忆替代指数衰减,在长上下文性能上持续优于 SOTA SSM 基线,同时在短上下文上保持竞争力。FRAC is introduced, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory and consistently improves long-context performance over state-of-the-art SSM baselines while staying competitive on short-context.
本文提出 Loop Scaling Laws,是首个将 recurrence 和 sparsity 与模型规模、数据联合建模的 scaling law,为在算力和显存约束下设计 looped MoE 模型提供了原则性基础。Loop Scaling Laws are introduced, the first scaling law to jointly model recurrence and sparsity alongside model size and data, and provide a principled foundation for designing looped MoE models under compute and memory constraints.
基于 RoPE 的语言模型出现长上下文失效,根源在于 RoPE 在维持稳定 token 偏好与区分相近位置之间存在固有权衡。要判断应处理哪种弱点以及如何处理,需要更精确地刻画 RoPE 在不同上下文长度下训练后模型中的行为。我们通过允许 RoPE 各频率下 query-key 尺度不一致,弥补了先前理论的一个关键局限,使之与实际经验观测高度吻合。我们的理论使得上述脆弱性对单个注意力头与输入可测量,并量化了高频Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns well with practical empirical observations. Our theory makes both vulnerabilities measurable for individual heads and inputs, and quantifies how high-freque
结果表明,有效的拒绝能够保留任务结构,同时限制与参考策略的耦合,并且较小的冻结模型可以低成本地提供此类拒绝能力。The results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.
本文提出 MemFold,根据所支持的行为来优化固定预算的软记忆;在 PersonaMem-32K 和 PersonaMem-128K 上取得了作者所测得的最高准确率,且在更长历史长度下优势进一步扩大。This work presents MemFold, which optimizes a fixed-budget soft memory by the behavior it supports, and attains the highest accuracy the authors measure on PersonaMem-32K and PersonaMem-128K, with margins that widen at the longer history length.
提出 FOVEATED,一种即插即用框架,通过随机偏移分配给前文上下文 key 的 Rotary Position Embedding 位置来构建每个句子的聚焦视图,并从理论上分析了 FOVEATED 如何抵消上下文引发的难度低估。FOVEATED, a plug-and-play framework that constructs focused views of each sentence by randomly shifting the Rotary Position Embedding positions assigned to the keys of its preceding context, is proposed and theoretically analyzed how FOVEATED counteracts context-induced difficulty underestimation is analyzed.
Extender 在 199M–924M 参数规模下匹配 Transformer 在短上下文(CORE)任务上的准确率,并在 924M 参数下于长上下文(RULER)任务上超越 Transformer。The Extender matches Transformer accuracy on short-context (CORE) tasks at 199M-924M parameters, and exceeds Transformer accuracy on long-context (RULER) workloads, again at 924M parameters.
提出 Behavior-Preserving KV Cache Compression,一种无需训练的框架,通过估计候选条目被移除后压缩缓存所诱导的 logits,并利用驱逐前的前向统计量评估其与完整缓存下一 token 分布之间的 KL 散度来为候选驱逐打分。This work proposes Behavior-Preserving KV Cache Compression, a training-free framework that scores candidate evictions by estimating the compressed-cache logits induced by their removal and evaluating the resulting KL to the full-cache next-token distribution using pre-eviction forward statistics.
结果表明,基于查询的循环记忆组合能够提升长上下文建模能力,并在超出训练上下文范围之后依然有效,同时每个 chunk 仅使用紧凑的仿射摘要。Results demonstrate that query-dependent composition of recurrent memory improves long-context modeling and remains effective beyond the training context while using compact per-chunk affine summaries.