研究库 论文知识库
Papers · organized/paper_cards

论文

187 张论文卡片 · LLM 基础设施

开放获取 全部 绿色 · 1640
6. LLM Research Papers: The 2026 List (Jan–May) — Sebastian Raschka
6. LLM Research Papers:2026 清单(1月–5月) — Sebastian Raschka
arXiv:2604.12374 LLM 基础设施 方法 OA · 绿色 被引 37 · S2

Nemotron 3 Super 是 Nemotron 3 系列中首个采用 NVFP4 进行预训练的模型,借助 LatentMoE(一种同时优化精度 per FLOP 与精度 per parameter 的新型 Mixture-of-Experts 架构),并集成 MTP 层以通过原生 speculative decoding 加速推理。Nemotron 3 Super is the first model in the Nemotron 3 family to be pre-trained in NVFP4, leverage LatentMoE, a new Mixture-of-Experts architecture that optimizes for both accuracy per FLOP and accuracy per parameter, and include MTP layers for inference acceleration through native speculative decoding.

Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You
无梯度自适应:仿射统计迁移及其证书能告诉你什么
arXiv:2609.00374 LLM 基础设施 应用落地 被引 0 · S2

本文提出 CASTER,一种无梯度方法:将源类统计量存储在判别子空间中,从目标批次矩估计一个类共享的仿射变换,并在分类前解析地将源类分布迁移到目标域,使其成为面向冻结模型部署的轻量适配机制。CASTER is introduced, a gradient-free method that stores source class statistics in a discriminative subspace, estimates a class-shared affine transformation from target-batch moments, and analytically transports the source class distributions before classification, which positions CASTER as a lightweight adaptation mechanism for frozen-model deployment.

CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing
CRISP:面向结构-质量驱动的悬崖感知输入自适应稀疏预填充
arXiv:2609.01925 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文用 C_struct 替代 Jensen-Shannon Divergence 路由:C_struct 是一种结构化代理,通过度量 Vertical-Slash 兼容位置上的 mass 来复现 JSD 的路由决策,同时消除池化 matmul 与后续 KL 散度开销。This work replaces the Jensen-Shannon Divergence routing with C_struct, a structural proxy that measures mass at Vertical-Slash compatible positions and reproduces JSD's routing decisions while eliminating both the pooled matmul and subsequent KL divergence overhead.

Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration
改变置信度,而非改变预测:面向后验校准的预测保持型修复
arXiv:2609.01072 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 CORD,第一个 post-fit adapter,通过从 calibrator 拟合中移除 preservation constraint 来修复完整校准概率向量,从而实现精确的预测保持,并将原始决策的精确恢复交由后续输出修复完成。CORD is proposed, the first post-fit adapter to impose exact prediction preservation by repairing the full calibrated probability vector by removing the preservation constraint from calibrator fitting and assigns exact recovery of the original decision to subsequent output repair.

Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance
优化中的渗流动力学:方差级联与离散尺度不变性
arXiv:2609.02373 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

通过将随机梯度流建模为渗流过程来研究随机梯度下降的动态,其中嵌套的架构对称性迫使子网络以离散块的形式合并,而非通过单边附着。The dynamics of Stochastic Gradient Descent is studied by modeling the stochastic gradient flow as a percolation process, in which nested architectural symmetries force subnetworks to merge in discrete blocks rather than by single-edge attachment.

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
锁于入口,开于内里:RLVR 收窄解空间的位置
arXiv:2608.29188 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

表面 prompting 未能恢复多样性,而针对入口的干预则成功:使用早期 checkpoint 的 late-layer parameter interpolation 在不损失 pass@1 的情况下将解的覆盖度提高了 37%。While surface prompting fails to recover diversity, entrance-targeted interventions succeed: late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1 and late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1.

4️⃣ Tangram · 多轮对话非均匀KV Cache — arXiv:2606.06302(⭐⭐⭐⭐ 新鲜 arXiv)
arXiv:2606.06302 LLM 基础设施 方法 OA · 绿色 被引 2 · S2

Tangram 是一种 serving 框架,将先前系统动态处理的内容静态解析,可作为现有 non-uniform 压缩方法的即插即用底座,在匹配其精度的同时,端到端吞吐量较 full-KV 基线最高提升 2.6×。Tangram is a serving framework that statically resolves what prior systems handle dynamically, and serves as a drop-in substrate for existing non-uniform compression methods, matching their accuracy while improving end-to-end throughput by up to $2.6\times over the full-KV baseline.

When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference
当量化破坏记忆时:低精度时序推理中的循环状态写回
arXiv:2609.04490 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

循环状态写回被确立为低精度循环动态的关键决定因素,state-storage interface 被识别为量化循环推理的核心设计考量。R recurrent-state write-back is established as a key determinant of low-precision recurrent dynamics and the state-storage interface is identified as a central design consideration for quantized recurrent inference.

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
BeaconKV:由 Beacon 查询引导的 KV Cache 压缩,用于高效大型推理模型推理
arXiv:2609.04971 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 BeaconKV,一种免训练的 KV cache 压缩方法,通过为每个全局查询簇维护紧凑代表性 beacon query 来预测哪些 KV 对将被重访,无需存储完整查询历史。BeaconKV is proposed, a training-free KV cache compression method that maintains beacon queries, compact representatives for each global query cluster, to anticipate which KV pairs will be revisited without storing the entire query history.

Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise
Encoded Early, Used Late:Transformer 在何处开始作用于所推断的对话伙伴专业度
arXiv:2609.07139 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

发现合作者专业能力在模型早期层最易被解码,到网络中点前降至接近随机水平,并基于合成语料以单个模型作为初步验证。It is found that partner expertise is most decodable in the early layers and falls to near chance before the midpoint of the network, and one model on a synthetic corpus is used as an initial demonstration.

Cadence: Error-Bounded Lossy Compression of Demand Time Series with a Time-Series Foundation Model
Cadence:基于时序基础模型的需求时序误差有界有损压缩
arXiv:2609.06008 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

我们提出 Cadence,一种面向数值时序的误差有界有损压缩器,将 3.3 亿参数的时序基础模型(Google TimesFM-3)与自适应算术编码器相结合,保证每个样本满足 |x_t - x̂_t| ≤ τ。一项负面结论限定了设计空间:在无损编码场景下,基础模型毫无价值,因为节省的比特数仅与预测器精度呈对数关系 Δb = log_2(MAE_old / MAE_new)。因此 TimesFM-3 相对 32 阶线性预测器 1.51 倍的精度优势,在 20.28 比特中仅换取 0.60 比特,中位数增益仅 +0.03%。误差有界编码仅在一点上突破了这一限制。We present Cadence, an error-bounded lossy compressor for numeric time series pairing a 330M-parameter time-series foundation model (Google TimesFM-3) with an adaptive arithmetic coder, guaranteeing |x_t-x_t|leτ on every sample. One negative result constrains the design space: for lossless coding a foundation model is worth nothing, because bits saved are logarithmic in predictor accuracy, Δb=log_2(MAE_{old}/MAE_{new}). So the 1.51times advantage TimesFM-3 holds over a 32-tap linear predictor buys 0.60 bits of 20.28, a median gain of +0.03%. Error-bounded coding escapes this at one point: once

3️⃣ Speculative Decoding 延迟可解释模型 — arXiv:2605.15051(⭐⭐⭐⭐ 调优必读)
arXiv:2605.15051 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文为 LLM serving 中的 SD(投机解码)提出了一种简单且可解释的 latency 模型,能够准确刻画实际观测到的 latency,解释为何加速比常随服务器负载上升而下降,并系统刻画了 draft length、acceptance rate 以及 verifier 与 drafter 规模在不同 serving 条件下对 latency 的影响。A simple and interpretable latency model for SD in LLM serving is developed that accurately describes observed latency, explains why speedups often diminish as server load increases, and characterizes how draft length, acceptance rate, and verifier-drafter size shape latency across serving conditions are characterized.

3. Tutti:让 SSD 后备 KV Cache 成为长上下文生产方案
arXiv:2605.03375 LLM 基础设施 方法 OA · 绿色 被引 6 · S2

Tutti 是一种高效的 SSD-backed KV caching 方案,将 CPU 从 HBM 与 SSD 之间的关键数据与 I/O 控制路径中彻底移除,在提供近乎无限容量的同时,实现了与 DRAM-backed LMCache 几乎相当的 inference 性能。Tutti is an efficient SSD-backed KV caching solution that eliminates CPU intervention from the critical data and I/O control paths between HBM and SSDs, and achieves nearly the same inference performance as DRAM-backed LMCache, while providing almost infinite capacity.

KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints
KVShareArena:跨上下文与模型 checkpoint 的 KV cache 复用
arXiv:2609.10266 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

揭示了复用造成的质量损失以及有效修复方法均取决于 LLM(即便在两个 8B 模型之间也是如此),并提供添加新方法的通用接口与交互式 leaderboardBoth the quality loss from reuse and which repairs help depend on the LLM, even between two 8B models are introduced, as well as a common interface for adding new methods and an interactive leaderboard.

The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements
稀疏性的代价:使用稀疏及稀疏化测量的稀疏恢复充分条件
arXiv:2509.01809 LLM 基础设施 方法 被引 0 · S2

证明:对于任意固定的目标误差水平 $\delta$ 与任意松弛量 $\varepsilon>0$,数量级为 $p/\psi^2$ 的样本量足以在任意小的 $\psi$ 下完成支撑恢复。It is proved that, for every fixed target error level $\delta$ and every slack $\varepsilon>0$, a sample size of order $p/\psi^2$ is sufficient for support recovery for arbitrarily small $\psi$.

3. Flow-Controlled Scheduling for LLM Inference(arXiv 2604.11001)
Flow-Controlled Scheduling for LLM Inference(arXiv 2604.11001)
arXiv:2604.11001 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出了一种简单的 flow-control 框架,通过控制 prompt 加入 LLM 活跃集合的速率,实现更高的 token 与 request 吞吐量、更低的平均与尾部 latency,以及更稳定的 KV cache 利用率。A simple flow-control framework is proposed that controls the rate at which prompts join the active set in large language models and achieves higher token and request throughput, lower average and tail latency, and more stable KV cache utilization.

Negative Self-Distillation: Learning to Reason by Avoiding Flaws
负自蒸馏:通过规避缺陷来学习推理
arXiv:2609.11699 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

提出 Negative Self-Distillation(NSD),一个通过偏离有缺陷的推理而非模仿特权解来优化 LLM 的新框架,并一致优于 OPSD 及其他无标签、自举式强化学习(RL)基线。Negative Self-Distillation (NSD) is introduced, a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions, and consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.

Beyond Solver Verdicts: Generative Reward Models for Autoformalization
超越求解器判定:面向自动形式化的生成式奖励模型
arXiv:2609.11085 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Generative Verification(GenV),通过复用语言模型的原生词表空间,将离线 Z3 等价性预言机蒸馏为无参考、连续参考等价分数,并从理论上证明基于结构、仅判决的验证启发式在这些欺骗性合法轨迹上的检测能力在数学上有界于随机水平。Generative Verification (GenV) is introduced, which distills an offline Z3-equivalence oracle into a reference-free, continuous reference-equivalence score by repurposing the language model's native vocabulary space and theoretically proves that structural, verdict-only verification heuristics are mathematically bounded to chance-level detection on these deceptively valid traces.

Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning
构建多语言桥梁:数据混合作为语言内推理泛化的支柱
arXiv:2609.10445 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

表明将 L2 推理泛化到未见语言的关键路径在于更广泛的语言覆盖、现成可用的多语言非推理数据,以及足够强的英语推理骨干,表明推理是一种与语言无关的行为,可通过精心数据混合在类型多样的语言间迁移。It is shown the path to generalizing L2 reasoning to held-out languages goes through broader language coverage, readily available multilingual non-reasoning data, and a sufficient English reasoning backbone, indicating that reasoning is a language-agnostic behavior that can be transferred across typologically diverse languages through careful data mixing.

Adaptive Bridge: A Proxy-Based Decoupling Layer for Mitigating DDS Backpressure in ROS 2
Adaptive Bridge:基于代理的解耦层以缓解 ROS 2 中 DDS 反压
arXiv:2608.15380 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,在评估 harness 中使用 Adaptive Bridge 可将关键的订阅者尾端 p95 延迟从最高 15 s 降至 1.55 ms,覆盖所有损伤严重度,同时保留发布者配置的吞吐量。The results show that using the Adaptive Bridge in the evaluation harness reduces the critical subscriber tail p95 latency from up to 15 s to 1.55 ms across all impairment severities while preserving the publisher's configured throughput.

Competence-Gated Pooling of Language Models and Priors for Event Forecasting
Competence-Gated Pooling of Language Models and Priors for Event Forecasting
arXiv:2609.12101 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出了一个能力门控(competence gate),根据已解决的结果估计领域级的源权重,将不确定的估计向全局权重收缩,并重新校准合并后的预测,为基于已测边际价值的选择性模型使用提供了实用方案。A competence gate is introduced that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast and provides a practical approach for selective model use based on measured marginal value.

2️⃣ Cats · 边缘推理的自投机级联验证 — arXiv:2605.11186(⭐⭐⭐⭐ 边缘推理重点)
arXiv:2605.11186 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 CATS,一种 self-speculative decoding 框架,在内存受限设备上结合 memory budget 与 parameter offloading pattern 进行级联式 verify 与 correction,在设备峰值显存与单独运行 target model 相当的前提下,最大化 token acceptance rate 与端到端加速比。CATS, a self-speculative decoding framework that conducts cascaded verification and correction based on the memory budget and parameter offloading patterns on memory-limited devices, is proposed, which maximizes token acceptance rate and end-to-end speedup while keeping the peak memory footprint on the device equal to that of the target model alone.

How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus
无损推测解码真的无损吗?数值精度在 Orthrus 中的作用
arXiv:2609.15504 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

展示并论证了算法无损性与其在有限精度算术下实现之间存在差距,主张应从精确生成轨迹与下游任务性能两个层面评估无损投机解码A gap between algorithmic losslessness and its implementation under finite-precision arithmetic is demonstrated and motivated, to motivate evaluating lossless speculative decoding at the level of exact generation trajectories as well as downstream task performance.

Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction
Grouped Value Attention:通过按需 Key 重建实现高效 KV 缓存
arXiv:2609.13285 LLM 基础设施 观点 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出了分组值注意力(Grouped Value Attention, GVA),通过存储分组值并在推理时利用可学习的线性映射重构内容键,该映射可被吸收进 query 中,从而在预期解码路径中无需显式物化内容键。This work introduces Grouped Value Attention (GVA), which stores grouped values and reconstructs content keys with a learned linear map at inference, which can be absorbed into the query, eliminating the need to materialize content keys in the intended decode path.

2. DualPath:打破 Agentic LLM 推理的存储带宽瓶颈
arXiv:2602.21548 LLM 基础设施 方法 Open MIND OA · 绿色 被引 20 · S2

本文提出 DualPath,一种 inference 系统,通过引入 dual-path KV-Cache loading 打破瓶颈,并实现一条新的 storage-to-decode 路径:KV-Cache 先加载到 decode engine,再通过 compute network 上的 RDMA 高效转发至 prefill engine。DualPath is presented, an inference system that breaks this bottleneck by introducing dual-path KV-Cache loading and enables a novel storage-to-decode path, in which the KV-Cache is loaded into decoding engines and then efficiently transferred to prefill engines via RDMA over the compute network.

RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs
RelateAnything: 任意输入的实时开放词汇关系预测
arXiv:2609.12552 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 RelateAnything,一个 53M 参数的模型,输入一张图像和来自任意来源的区域,针对推理时以字符串形式提供的谓词词汇返回带分数的关系。This work presents RelateAnything, a 53M-parameter model taking an image and regions from any source and returning scored relations over a predicate vocabulary supplied at inference as strings, a 53M-parameter model taking an image and regions from any source and returning scored relations over a predicate vocabulary supplied at inference as strings.

1️⃣2️⃣ arXiv · Cloud-native and Distributed Systems for LLM:研究路线图 ⭐⭐⭐ 学术综述
arXiv · Cloud-native and Distributed Systems for LLM:研究路线图 ⭐⭐⭐ 学术综述
arXiv:2604.17227 LLM 基础设施 综述 OA · 绿色 被引 4 · S2

本文探讨了 cloud platform 与 distributed system 在支撑 LLM 可扩展性、效率与优化方面的作用,涵盖数据管理、资源优化,以及对 microservices、autoscaling 与 hybrid cloud-edge 方案的需求。The role of cloud platforms and distributed systems in supporting the scalability, efficiency, and optimization of LLMs is explored, including data management, resource optimization, and the need for microservices, autoscaling, and hybrid cloud-edge solutions.

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
重新思考 PPO 中的 Critic 学习:理解与缓解 Value Flattening
arXiv:2609.18708 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

本文识别出 Value Flattening 是标准 PPO 中 critic 学习的一种重要但被忽视的失效模式,并提出一种简单稀疏监督策略可以缓解该问题;引入 SParse Proximal Policy Optimization,在每个响应中仅对少数间隔良好的状态施加 value loss,以同时缓解两种效应。Value Flattening is identified as an important yet overlooked failure mode of critic learning in standard PPO and a simple sparse supervision strategy can mitigate it; SParse Proximal Policy Optimization is introduced, which applies the value loss to only a few well-separated states in each response to mitigate both effects.

The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
内存墙的另一半:通过训练路由预测从 SSD 服务 35B MoE
arXiv:2609.18063 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Edge0,一个流式 MoE 推理引擎,通过 prerouter 缩小差距:每层 head 提前一个 token 预测下一层的 routing,并将该预测直接作为 routing 使用,使分阶段 expert 集合等于路由集合,无任何丢弃。Edge0 is presented, a streaming MoE inference engine that closes the gap with a prerouter: a per-layer head predicts the next layer's routing one token ahead, and the prediction is consumed as the routing itself, so the staged expert set equals the routed set and nothing is dropped.

Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches
Fathom:面向卸载 KV 缓存稀疏解码的每查询读取深度
arXiv:2609.17652 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Fathom,一种 key scan,其中每个查询决定读取每个 key 通道的比特数;在真实的 coding agent 会话中,Fathom 以 92 bit 达到最准确的 136 bit scan 的步骤一致性。Fathom is presented, a key scan in which each query decides how many bits of each key channel to read, and on real coding-agent sessions Fathom reaches the step agreement of the most accurate 136-bit scan at 92 bits.

Assessing nnU-Net Generalization across Brain Tumor Populations in BraTS-GoAT 2026
在 BraTS-GoAT 2026 中评估 nnU-Net 跨脑肿瘤人群的泛化能力

BraTS-GoAT 在异质人群上使用传统 3D nnU-Net 对 1,351 个标注案例进行肿瘤分割评估,采用五折交叉验证,每折 1,000 epochs,并应用 test-time mirroring。BraTS-GoAT evaluates tumor segmentation across heterogeneous populations using a conventional 3D nnU-Net on 1,351 labeled cases using five-fold cross-validation and 1,000 epochs per fold and applied test-time mirroring.

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
DeepSeek-V4.1-Flash:突破 KV Cache 压缩的极限
arXiv:2609.19969 LLM 基础设施 应用落地 OA · 绿色 被引 43 · S2

推出 DeepSeek-V4.1-Flash 模型,这是一个具有 552B 骨干参数、支持最长一百万 token 上下文的多模态 Mixture-of-Experts 模型,显著提升了 agent 工作负载的成本效率,并突破了 KV cache 压缩的极限。The DeepSeek-V4.1-Flash model, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts of up to one million tokens, is introduced, substantially improving cost efficiency for agentic workloads and pushing the limits of KV cache compression.

WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing
WeVisDoc:从覆盖到能力的鲁棒端到端文档解析
arXiv:2609.20423 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 WeVisDoc,一个面向鲁棒端到端文档解析的两阶段数据驱动框架,通过异构数据与保持结构的退化合成来扩展语义、结构与外观覆盖度,并指导有针对性的数据构建与目标 token 预算的重分配。WeVisDoc is presented, a two-stage data-centric framework for robust end-to-end document parsing that broadens semantic, structural, and appearance coverage through heterogeneous data and structure-preserving degradation synthesis and guides targeted data construction and reallocation of the target-token budget.

What Does Privileged Information Add to On-Policy Self-Distillation?
特权信息为 On-Policy 自蒸馏带来了什么?
arXiv:2609.20612 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

构建了 AMPLE-Math,一个包含 5,319 道数学问题、每题对应六种共享同一答案的推理视角的可复用套件,并通过对比表明 OPSD 能借助直答推理与带思维链推理所共享的参数,提升对已有推理能力的调用效率。AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, is constructed and AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, is compared, suggesting that OPSD can improve access to existing reasoning capabilities through parameters shared by direct-response and thinking-enabled inference.

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
当 EOS token 不一致时:理解 On-Policy 蒸馏中的长度膨胀
arXiv:2609.20511 LLM 基础设施 方法 OA · 绿色 被引 2 · S2

我们研究 on-policy 蒸馏 (OPD) 中的长度膨胀现象,即学生回答会变得过长,甚至耗尽生成预算。我们发现基础学生模型与训练后教师模型之间的终止 token 不匹配是该行为的重要来源。在 Qwen3、Llama 和 Gemma 上,两个模型可能将停止概率分配到不同的 EOS token 上,即便它们声明的停止集合相同。这种不匹配会抑制学生偏好的终止动作,同时无法可靠地传递教师偏好的替代动作。我们证明对齐 t...We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify termination-token mismatch between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama, and Gemma, the two models can place their stopping probability on different EOS tokens, even when their declared stopping sets are identical. This mismatch can suppress the student's preferred termination action without reliably transferring the teacher-preferred alternative. We show that aligning t

Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
样本数量远远不够:候选生成策略决定 LLM 测试时扩展的能耗与性能
arXiv:2609.19499 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,仅靠候选数量不足以刻画多候选 test-time scaling 的系统成本;评测应同时报告候选数量与准确率,以及生成调度和 GPU 层级的系统指标。The results show that candidate count alone is not enough to describe the systems cost of multi-candidate test-time scaling and Evaluations should report not only candidate count and accuracy, but also generation schedule and GPU-level systems metrics.