政治话语日益呈现更强的敌对感是普遍观感,但稳健证据仍稀缺。本文研究 1997 至 2026 年丹麦议会的归责行为,结合专门构建的分类器 BlameBERT(F1: 0.80)与多层统计建模。该分类器采用面向低至中等资源语言的标注高效流程。结果显示出一条香蕉形轨迹:归责水平约在 2016 年前下降,随后在近年(2019–2026)进入显著且持续的上升阶段。执政地位显著Political discourse is widely perceived to be growing more hostile, yet robust evidence remains scarce. This study examines blame attribution in the Danish Parliament from 1997 to 2026, combining a purpose-built classifier, BlameBERT (F1: 0.80), with multilevel statistical modeling. The classifier is constructed using an annotation-efficient pipeline for blame attribution in low-to-mid resource languages. The results reveal a banana-shaped trajectory, with blame declining until around 2016 before entering a significant and sustained increase in recent years (2019-2026). Government status consi
论文
177 张论文卡片 · LLM 基础设施 · OA 绿色
本文展示了在一对架构匹配的 Qwen3.5 4B-to-9B 兄弟模型之间实现有用的持久 hybrid-state transfer,是首次在大小不同的混合语言模型之间完成持久循环推理状态的跨模型交接,且无需 target prefix replay。This work demonstrates useful persistent hybrid-state transfer across one architecture-matched Qwen3.5 4B-to-9B sibling pair, the first demonstrated cross-model handoff of persistent recurrent inference state between differently sized hybrid language models without target prefix replay.
TritonForge 是一个面向自动化 Triton kernel 优化的 profiling 引导框架,融合 kernel 分析、运行时 profiling 与迭代式代码转换以简化优化流程,并为自动化 GPU 性能优化领域的未来研究奠定基础。TritonForge, a profiling-guided framework for automated Triton kernel optimization that integrates kernel analysis, runtime profiling, and iterative code transformation to streamline the optimization process and provides a foundation for future research in automated GPU performance optimization.
本文探讨容器化、微服务、动态调度等 Cloud Native 技术如何从根本上提升 LLM 推理服务,并展示 Cloud Native 系统在高需求场景下实现更高效资源分配、降低延迟与提升吞吐的能力。This article explores how Cloud Native technologies, such as containerization, microservices, and dynamic scheduling, can fundamentally improve LLM inference serving and demonstrates how a Cloud Native system enables more efficient resource allocation, reduces latency, and enhances throughput in high-demand scenarios.
本文提出一种方法,通过逐层剔除后 WER(Word Error Rate)的变化对编码器层进行排序,并给出剪枝后的模型——即一个层数更少的更浅编码器。This work presents an approach that ranks encoder layers by the leave-one-layer-out change in Word Error Rate (WER), and presents the pruned model, which is simply a more shallow encoder with fewer layers.
本文论证线性表征假设不是单一假设,而是一组通过表征等价性加以区分的命题族;该框架厘清了假设如何随度量、读取点与分析阶段变化,并被用于审计常见表征量与近期可解释性分析。It is argued that the Linear Representation Hypothesis is not one hypothesis but a family of claims distinguished by representation equivalence, and this framework clarifies how assumptions can change across metrics, reading points, and analysis stages, and is used to audit common representation quantities and recent interpretability analyses.
本文首次对 vLLM 启动延迟进行了详细的性能表征,并构建了一个轻量级分析模型,能够针对给定硬件配置准确预测 vLLM 的启动延迟,为大规模推理环境中的资源规划提供了可操作的指导。This paper presents the first detailed performance characterization of vLLM startup latency and develops a lightweight analytical model that accurately predicts vLLM's startup latency for a given hardware configuration, providing actionable guidance for resource planning in large-scale inference environments.
本文提出 Neural Spectral Capacity (NSC),一种以每个权重矩阵的奇异值谱为依据的闭式标量,在七类 Transformer 与 CNN 系列上的排序效果优于 #Params、#FLOPs 及代表性免训练代理指标。This work proposes Neural Spectral Capacity (NSC), a closed-form scalar grounded in the singular-value spectrum of each weight matrix, which outperforms #Params, #FLOPs, and representative training-free proxies in ranking across seven Transformer and CNN families.
本文提出 PISA,一种采用金字塔式 Top-K 选择策略的 block-sparse attention 机制,并为训练与推理开发了硬件感知的 Triton kernel,将层级路由与 LogSumExp 评分融合,无需显式构造 query-key score 矩阵。PISA is proposed, a block-sparse attention mechanism that employs a pyramid Top-$K selection strategy, and develops hardware-aware Triton kernels for both training and inference, fusing hierarchical routing and LogSumExp scoring without materializing the query-key score matrix.
本文提出 PsPLUG,一个轻量级插件式模块,在刻画目标风格后学习用户专属的残差,能更好地保留用户偏好,并对个性化与风格遵循之间的平衡提供更精确的控制。PsPLUG is proposed, a lightweight plug-in that learns a user-specific residual after accounting for the requested style, and better preserves user preferences while providing precise control over the balance between personalization and style adherence.
本文提出首个通过 continuous depth batching (CDB) 实现 depth-adaptive looped LM 的高效方法:在 loop 步骤之间重组 batch,动态调度架构中的 looped 与非 looped 部分,管理 looped KV-caching,并提前预测将退出 loop 的 token,以便异步准备 batch。This work introduces the first efficient method for depth-adaptive looped LMs via continuous depth batching (CDB), which forms new batches between loop steps, and dynamically schedules looped and non-looped parts of the architecture, manages looped KV-caching, and predicts which tokens will exit the loop in advance so it can prepare batches asynchronously.
使用一种分层旋转位置编码(hRoPE),将段落、句子和 token 索引表示为独立通道,保持 token 序列固定,对段落坐标 $p_1$ 进行干预,并使用 token 距离精确估计器测量跨段落注意力。A hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate $p_1$, and measure cross-paragraph attention with a token-distance-exact estimator is used.
Boxoffice 是一个程序化生成评测数据集的工具,用以覆盖具有挑战性的 KV cache 复用模式;研究表明现有数据集并未展现出充分评估此类技术所需的复用动态特性。Boxoffice, a tool that programmatically generates evaluation datasets that exercise challenging KV cache reuse patterns, is introduced and it is shown that existing datasets do not exhibit the reuse dynamics needed to thoroughly evaluate such techniques.
Intent2Tc 是一个由语言模型驱动的闭环框架,将业务级流量整形意图转换为声明式子意图,进而生成经过验证的可执行 Linux traffic control (tc) 配置,并展示了该框架的实际适用性。Intent2Tc is presented, a closed-loop language-model-driven framework that translates business-level traffic-shaping intents into declarative sub-intents and subsequently into validated, executable Linux traffic control (tc) configurations and demonstrates the practical applicability of the proposed framework.
本文提出 softmax reparameterization,一种训练后方法,在量化前搜索功能等价的输出头,并展示了在总体 logit 误差增大的情况下保真度仍可提升。This work introduces softmax reparameterization, a post-training method that searches over functionally equivalent output heads before quantization and shows how fidelity can improve despite greater total logit error.
评估 vLLM 在分离式 serving 下的精度,并通过训练后量化在最高 2.8T 参数的模型上进一步验证共享权重格式的分离式 serving。Assessment of accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters are evaluated.
NetKV,一种使用该 oracle 信息的 O(|D|) 每请求贪心策略,其层级排序被证明对过时遥测数据具有鲁棒性;并证明随着上下文长度增长,忽略网络项会使仅缓存感知的调度任意次优。NetKV, the O(|D|) per-request greedy that consumes this oracle, has tier rankings that are provably robust to stale telemetry, and it is proved that ignoring the network term renders cache-aware-only scheduling arbitrarily suboptimal as context length grows.
受 Apache Spark 等分布式计算框架启发,DISCO 将长上下文切分到一组专职 Worker LLM 上,进行并行、局部的 grounding,建立了鲁棒长上下文推理的高效范式。Inspired by distributed computing frameworks like Apache Spark, DISCO partitions long context across a fleet of Worker LLMs dedicated exclusively to parallel, localized grounding, establishing a highly efficient paradigm for robust long-context inference.
本工作提出 LMCACHE,首个也是目前最高效的开源 KV 缓存方案,可将现代 LLM 引擎生成的 KV 缓存从 GPU 显存中提取并存储,并跨引擎和查询共享。This work presents LMCACHE, the first and so far the most efficient open-source KV caching solution, which extracts and stores KV caches generated by modern LLM engines out of the GPU memory and shares them across engines and queries.
提出 G²PTQ,一个具有广义梯度补偿的统一 PTQ 框架,在全局监督的分块优化目标下融合一阶与二阶信息,能更好地对齐全精度模型,性能优于 SOTA 基线。G$^2$PTQ is presented, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective and enables better alignment with the full-precision model, outperforming state-of-the-art baselines.
通道级 history-state 注入结合 timestep 条件化构成了一种低成本且更有效的设计,能够在更长 unrolling 下更好地保留知识,同时提升在不同推理预算下的鲁棒性。It is shown that channel-wise history-state injection combined with timestep conditioning offers a low-cost and more effective design, better preserving knowledge under extended unrolling while improving robustness across inference budgets.
本文揭示使用 chunked KV-cache 压缩的模型中存在一种系统性不对称:相同信息在一个阶段易于检索,在另一阶段却难以获取,暴露出平均 benchmark 分数可能掩盖的周期性薄弱点。A systematic asymmetry in models using chunked KV-cache compression is uncovered: the same information can be easy to retrieve at one phase and difficult at another, revealing periodic weak spots that average benchmark scores can conceal.
本文提出 FRAC,一种源自分数阶动力学的选择性 SSM 架构,用幂律长记忆替代指数衰减,在长上下文性能上持续优于 SOTA SSM 基线,同时在短上下文上保持竞争力。FRAC is introduced, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory and consistently improves long-context performance over state-of-the-art SSM baselines while staying competitive on short-context.
本文提出 Loop Scaling Laws,是首个将 recurrence 和 sparsity 与模型规模、数据联合建模的 scaling law,为在算力和显存约束下设计 looped MoE 模型提供了原则性基础。Loop Scaling Laws are introduced, the first scaling law to jointly model recurrence and sparsity alongside model size and data, and provide a principled foundation for designing looped MoE models under compute and memory constraints.
基于 RoPE 的语言模型出现长上下文失效,根源在于 RoPE 在维持稳定 token 偏好与区分相近位置之间存在固有权衡。要判断应处理哪种弱点以及如何处理,需要更精确地刻画 RoPE 在不同上下文长度下训练后模型中的行为。我们通过允许 RoPE 各频率下 query-key 尺度不一致,弥补了先前理论的一个关键局限,使之与实际经验观测高度吻合。我们的理论使得上述脆弱性对单个注意力头与输入可测量,并量化了高频Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns well with practical empirical observations. Our theory makes both vulnerabilities measurable for individual heads and inputs, and quantifies how high-freque
结果表明,有效的拒绝能够保留任务结构,同时限制与参考策略的耦合,并且较小的冻结模型可以低成本地提供此类拒绝能力。The results suggest that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide them at low cost.
本文提出 MemFold,根据所支持的行为来优化固定预算的软记忆;在 PersonaMem-32K 和 PersonaMem-128K 上取得了作者所测得的最高准确率,且在更长历史长度下优势进一步扩大。This work presents MemFold, which optimizes a fixed-budget soft memory by the behavior it supports, and attains the highest accuracy the authors measure on PersonaMem-32K and PersonaMem-128K, with margins that widen at the longer history length.
提出 FOVEATED,一种即插即用框架,通过随机偏移分配给前文上下文 key 的 Rotary Position Embedding 位置来构建每个句子的聚焦视图,并从理论上分析了 FOVEATED 如何抵消上下文引发的难度低估。FOVEATED, a plug-and-play framework that constructs focused views of each sentence by randomly shifting the Rotary Position Embedding positions assigned to the keys of its preceding context, is proposed and theoretically analyzed how FOVEATED counteracts context-induced difficulty underestimation is analyzed.
提出 Behavior-Preserving KV Cache Compression,一种无需训练的框架,通过估计候选条目被移除后压缩缓存所诱导的 logits,并利用驱逐前的前向统计量评估其与完整缓存下一 token 分布之间的 KL 散度来为候选驱逐打分。This work proposes Behavior-Preserving KV Cache Compression, a training-free framework that scores candidate evictions by estimating the compressed-cache logits induced by their removal and evaluating the resulting KL to the full-cache next-token distribution using pre-eviction forward statistics.
结果表明,基于查询的循环记忆组合能够提升长上下文建模能力,并在超出训练上下文范围之后依然有效,同时每个 chunk 仅使用紧凑的仿射摘要。Results demonstrate that query-dependent composition of recurrent memory improves long-context modeling and remains effective beyond the training context while using compact per-chunk affine summaries.
提出基于流的自潜在推理(Flow-based Latent Reasoning, FLaRe),一份简洁方案,涵盖潜空间编码内容及其塑形方式、流训练位置、答案读取方式,以及最终在模型自身已验证思考上进行训练的阶段。Flow-based Latent Reasoning (FLaRe), a simple recipe covering what the latent space encodes and how to shape it, where to train the flow, how to read out the answer, and a final stage of training on the model's own verified thoughts, is presented.
这是一篇综述文章,围绕机器学习技术(特别是大语言模型)展开,旨在以对数值方法领域专家友好的方式阐释 LLM 的核心概念。This is a survey article that centers on machine learning techniques, with a particular focus on large language models, to clarify the core concepts behind Large Language Models in a manner accessible to specialists in numerical methods.
提出 KVpop,通过直接监督保留或丢弃决策来学习固定预算的 KV 淘汰策略,并引入基于延迟记忆的打分器,在已学习的淘汰方法中独有地延迟固定步数的打分,以利用近期上下文。KVpop is introduced, which learns a fixed-budget KV eviction policy by directly supervising the keep-or-drop decision, and introduces a delayed memory-based scorer that, uniquely among learned eviction methods, defers scoring for a fixed number of steps to exploit near-future context.
提出流形约束假设(Manifold Constraint Hypothesis, MCH):若自然表征集中于一个结构化的低维流形上,则干预应被约束在该流形内,从而在干预过程中更好地保留表征中编码的其他信息。The Manifold Constraint Hypothesis (MCH) is proposed: if natural representations concentrate on a structured, lower-dimensional manifold, then interventions should be constrained to that manifold and better preserve other information encoded in the representation during interventions.
提出 Simplified Sparse Attention——一种更简洁的稀疏注意力方案,无需任何架构改动,在检索增强生成中优于全注意力;并扩展为分层 gist-of-gist 变体,在高达 32× 的高压缩比下保持或提升精度,同时实现对数级解码复杂度。Simplified Sparse Attention is introduced, a simpler approach to sparse attention that requires no architectural changes that outperforms full attention in retrieval-augmented generation and extends to a hierarchical gist-of-gist variant that achieves log-linear decoding complexity while maintaining or improving accuracy at high compression ratios up to 32x.
提出 Erase-then-Delta Attention (EDA),一种将"在哪里擦除"与"在哪里写入"解耦的内存更新规则;研究表明循环记忆模型不仅应决定写入什么,还应决定擦除哪些陈旧信息以及擦除的位置。Erase-then-Delta Attention (EDA), a memory update rule that decouples where to erase from where to write, is proposed, suggesting that recurrent memory models should decide not only what to write, but also what stale information to erase and where.