研究库 论文知识库
Papers · organized/paper_cards

论文

143 张论文卡片 · LLM 基础设施 · 方法 · OA 绿色

开放获取 全部 绿色 · 1640
5. SwiftCache: Efficient LLM Serving for Multi-turn Conversations
SwiftCache:面向多轮对话的高效 LLM serving
arXiv:2606.16135 LLM 基础设施 方法 OA · 绿色 被引 2 · S2

SwiftCache 是一个协同推理系统,使异构模型可在同一服务器内共享未充分利用的 GPU 内存与 NVLink 带宽,支持跨模型通过 NVLink 共享 KV cache,避免使用慢速 PCIe 传输。SwiftCache is a collaborative inference system that enables heterogeneous models to share underutilized GPU memory and NVLink bandwidth within a server, allowing cross-model KV cache sharing over NVLink and avoiding slow PCIe transfers.

6️⃣ OScaR · 极端KV Cache量化 — arXiv:2605.19660(⭐⭐⭐ arXiv)
🔴 保留 · OScaR · 极端 KV cache 量化 — arXiv:2605.19660(⭐⭐⭐ arXiv)
arXiv:2605.19660 LLM 基础设施 方法 OA · 绿色 被引 4 · S2

提出 OScaR(Omni-Scaled Canalized Rotation),一个面向 X-LLMs 的精确且轻量的 KV cache 压缩框架,作为一个鲁棒、低复杂度、通用的框架确立了新的 Pareto 前沿。OScaR (Omni-Scaled Canalized Rotation), an accurate and lightweight KV cache compression framework for X-LLMs, is proposed, establishing it as a robust, low-complexity, and universal framework that defines a new Pareto front.

5️⃣ arXiv · Fluid-Guided在线调度 + WAIT策略(⭐⭐⭐⭐ 补充)
arXiv:2504.11320 LLM 基础设施 方法 OA · 绿色 被引 24 · S2

WAIT (Waiting for Accumulated Inference Threshold) 是一种基于阈值的准入规则,适用于已知输出长度;Nested WAIT 通过调控请求在 decode 阶段各分段间的推进方式,将该规则扩展到未知输出长度场景。WAIT (Waiting for Accumulated Inference Threshold), a threshold-based admission rule for known output lengths, and Nested WAIT, which extends the rule to unknown output lengths by regulating how requests advance across decode-stage segments are designed.

1️⃣ arXiv · AIConfigurator:多框架LLM推理配置自动优化(⭐⭐⭐⭐⭐ 必读)
arXiv:2601.06288 LLM 基础设施 方法 OA · 绿色 被引 15 · S2

本文提出 AIConfigurator,一个统一的性能建模系统,能够在不依赖 GPU profiling 的前提下进行快速、与框架无关的推理配置搜索;并提供一个抽象层,自动为目标后端解析最优启动参数,无缝集成到生产级编排系统中。AIConfigurator is presented, a unified performance-modeling system that enables rapid, framework-agnostic inference configuration search without requiring GPU-based profiling, and an abstraction layer that automatically resolves optimal launch parameters for the target backend, seamlessly integrating into production-grade orchestration systems.

论文信息
arXiv:2512.24601 LLM 基础设施 方法 Open MIND OA · 绿色 被引 90 · S2

研究发现 RLMs 能够成功处理超出模型上下文窗口长达两个数量级的输入,即便在较短 prompt 下,其质量也显著优于原生前沿 LLM 以及常见的长上下文与编程脚手架。It is found that RLMs can successfully process inputs up to two orders of magnitude beyond model context windows and, even for shorter prompts, dramatically outperform the quality of vanilla frontier LLMs and common long-context and coding scaffolds.

论文信息
arXiv:2606.11916 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出了一种实证方法,用于研究基于 GPU 的 LLM 推理服务系统中的 software aging 问题,并提供了一个可复现的框架,开辟了 software aging 与 software rejuvenation 与 LLM serving 交叉方向的研究。This paper proposes an empirical methodology to study software aging in GPU-based LLM serving systems and provides a reproducible framework that opens a research direction at the intersection of the software aging and rejuvenation and LLM serving communities.

条目E4:arXiv 2605.04595 — KV Cache 队列论理与稳定性分析
arXiv:2605.04595 LLM 基础设施 方法 OA · 绿色 被引 4 · S2

本文提出首个将计算与 GPU 显存约束显式纳入 LLM 推理分析的排队论框架,并推导了严格的稳定性与不稳定性条件,用以判定 LLM 推理服务能否在持续到达的请求下避免队列无界增长。This paper introduces the first queueing-theoretic framework that explicitly incorporates both computation and GPU memory constraints into the analysis of LLM inference, and derives rigorous stability and instability conditions that determine whether an LLM inference service can sustain incoming demand without unbounded queue growth.

条目 E-NF1:Albireo — 突破 Amdahl 定律的 LLM 推理张量并行调度
arXiv:2606.01927 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Albireo,一种并行推理系统:通过调度与 I/O 与计算及序列并行采样的重叠来压缩不可扩展部分,从而提高可达到的 TP 度 t,且不改变模型架构。Albireo is presented, a parallel inference system that raises the attainable TP degree t by shrinking the non-scalable portion via overlap of scheduling and I/O with compute and sequence-parallel sampling, without changing model architectures.

Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion
超越故事中心数据的可扩展创意写作:基于属性引导的题材扩展
arXiv:2608.13947 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

实验表明,在该数据上微调的模型不仅持续优于基座模型和面向写作的专门基线,还优于基于现有写作语料训练的模型,表明受控的题材扩展是稳健创意写作能力的关键驱动力Experiments demonstrate that models fine-tuned on the data consistently surpass not only base models and writing-specialized baselines, but also models trained on existing writing corpora, indicating that controlled genre expansion is a key driver of robust creative writing capability.

Chain-of-Experience for Continual LLM Improvement
Chain-of-Experience:基于经验链的 LLM 持续改进
arXiv:2608.18027 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文研究 LLM 在测试时从迭代经验中学习的机制,称之为 Chain-of-Experience (CoE):模型通过与自身或环境反馈的迭代交互积累经验痕迹,形成超越零样本推理的持续改进循环。This study studies how LLMs learn from iterative experience at test time, a setting the authors refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference.

Towards a Densing Law for User Representation Learning at Billion-Scale Capacity
迈向十亿级容量用户表示学习的稠密定律
arXiv:2608.23392 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 User Behavioral Densing Law,为大规模用户表示学习中的 tokenization 配置选择提供实用指导,并开发了 ALGN——一种自适应变长 tokenization 方法,可改善容量分配。The proposed User Behavioral Densing Law is proposed, providing practical guidance for tokenization configuration selection in large-scale user representation learning and ALGN, an adaptive variable-length tokenization method that improves capacity allocation, is developed.

7. Triton Attention Kernel 学术分析 (arXiv 2511.11581)
Triton Attention Kernel 学术分析 (arXiv 2511.11581)
arXiv:2511.11581 LLM 基础设施 方法 OA · 绿色 被引 4 · S2

本工作开发了一个 SOTA 的 paged attention kernel,完全基于领域特定即时编译语言 Triton 构建,在 NVIDIA 与 AMD GPU 上均达到 SOTA 性能。This work develops a state-of-the-art paged attention kernel that builds exclusively on the domain-specific just-in-time compiled language Triton to achieve state-of-the-art performance on both NVIDIA and AMD GPUs.

7. LLM 压缩:联合剪枝 + 混合精度 PTQ
arXiv:2606.07819 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出一种新颖的混合精度 PTQ 策略,直接最小化整个模型的全局误差传播,而非孤立地处理逐层误差,并开发了一种新颖的联合优化方法,在统一搜索空间中同时学习结构化剪枝决策与混合精度量化策略。This work proposes a novel mixed-precision PTQ strategy that directly minimizes global error propagation across the entire model, rather than isolating layer-wise errors, and develops a novel joint optimization approach that simultaneously learns structural pruning decisions and mixed-precision quantization policies within a unified search space.

Prefix Sliding for efficient test-time scaling
Prefix Sliding:面向高效 test-time scaling 的前缀滑动方法
arXiv:2608.26070 LLM 基础设施 方法 OA · 绿色 被引 2 · S2

提出 Prefix Sliding,在推理过程中丢弃不属于前缀或最近几千 token 窗口的 token,从而实现高效的长时程 test-time scaling。Prefix Sliding is proposed, which discards tokens during reasoning that are not part of the prefix or the window of the last few thousand tokens, allowing for efficient long-horizon test-time scaling.

Scaling Laws for Neural Language Models
神经语言模型的 Scaling Laws
arXiv:2001.08361 LLM 基础设施 方法 OA · 绿色 被引 9195 · S2

更大的模型显著更具样本效率,因此最优的算力高效训练方式是:在相对适中的数据量上训练非常大的模型,并在远未收敛时显著提前停止训练。Larger models are significantly more sample-efficient, such that optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence.

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes
CritICL:基于小语言模型失败模式的推理时弱到强泛化
arXiv:2608.27455 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

实验结果表明,CritICL 持续优于标准 in-context learning,并以显著更少的生成次数和更低的 token 成本,取得与 test-time scaling 方法相当或更优的性能。Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost.

Fast Weight Attention for Continual Learning
用于持续学习的快速权重注意力
arXiv:2608.27763 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该框架将循环序列模型中的 temporal alignment、plasticity、forgetting 与 bounded rehearsal 分离,并结合数值稳定的 positive-decay renormalization,使其在语言建模上保持竞争力,同时提升在可变位数加法任务上的长度外推能力。This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models, together with numerically stable positive-decay renormalization, to remain competitive in language modeling and improve length extrapolation on variable-digit addition.

GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models
GGSS:用于生成式视觉-语言模型推理时去偏的测地门控球面引导
arXiv:2608.25375 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

GGSS——Geodesic-Gated Spherical Steering——一种保范的干预方法,在单位超球面上发现反事实偏差子空间,沿 geodesic arc 引导视觉 token,并通过自适应门控聚焦于承载更强人口统计信号的 token。GGSS---Geodesic-Gated Spherical Steering---a norm-preserving intervention that discovers a counterfactual bias subspace on the unit hypersphere, steers visual tokens along geodesic arcs, and uses an adaptive gate to focus correction on tokens that carry stronger demographic signal.

Sliding-window beats linear attention
滑动窗口优于线性注意力
arXiv:2608.28444 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作表明带 sink 的 Sliding Window Attention(SWA)表现不逊于甚至优于后训练的 Linear Attention 模型,并建议改用 SWA 而非后训练线性模型。This work shows that Sliding Window Attention (SWA) with sinks performs as well or better than post-trained Linear Attention models, and recommends switching to SWA instead of post-training linear models.

Uncertainty-Aware End-to-End AI Weather Forecasting: Disentangling Observation and Model Contributions
感知不确定性的端到端 AI 天气预报:分离观测与模型的贡献
arXiv:2608.30795 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

端到端天气预报系统直接从原始地球观测数据生成高质量的全球网格与站点预报,仅以数值天气预报流水线(含数据同化)一小部分的成本取而代之。这类系统是确定性的,不输出不确定性结果。本文通过在每个组件上附加一种随机机制,将 Aardvark Weather 模型概率化:在观测编码器中加入学习到的、依赖输入的噪声,以捕捉源自观测系统的偶然不确定性;在处理器中使用 Monte Carlo dropout,以捕捉认知不确定性。End-to-end weather forecasting systems produce skillful global gridded and station forecasts directly from raw Earth observations, replacing the numerical weather prediction pipeline, including data assimilation, at a fraction of its cost. These systems are deterministic and issue no uncertainty. Here we render the Aardvark Weather model probabilistic by attaching one stochastic mechanism to each component: learned, input-dependent noise at the observation encoder, capturing aleatoric uncertainty inherited from the observing system, and Monte Carlo dropout in the processor, capturing epistemic

DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory
DiagEvo:通过分层错误记忆实现的诊断引导自我演化
arXiv:2609.00768 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

本文提出 DagEvo,通过利用 self-play 中求解器的失败历史来引导问题生成,无需外部任务资源,并表明混合生成、跨状态拼接的记忆状态更新以及双置信度过滤共同贡献了这些性能提升。DagEvo is introduced, which guides question generation using the solver's failure history from self-play, without external task resources, and shows that mixed generation, memory-state updates with cross-state stitching, and double-confidence filtering contribute to these gains.

6. LLM Research Papers: The 2026 List (Jan–May) — Sebastian Raschka
6. LLM Research Papers:2026 清单(1月–5月) — Sebastian Raschka
arXiv:2603.15569 LLM 基础设施 方法 OA · 绿色 被引 99 · S2

本工作借鉴线性模型的 state space model(SSM)视角,提出三项核心方法改进并组合形成更具表达力的递推结构:源自 SSM 离散化的递推式、用于更丰富状态追踪的复数值状态更新规则,以及在不增加 decode 延迟前提下提升模型性能的多输入多输出(MIMO)建模。This work introduces three core methodological improvements inspired by the state space model (SSM) viewpoint of linear models that combine to form a more expressive recurrence derived from SSM discretization, a complex-valued state update rule that enables richer state tracking, and a multi-input, multi-output (MIMO) formulation for better model performance without increasing decode latency.

6. LLM Research Papers: The 2026 List (Jan–May) — Sebastian Raschka
6. LLM Research Papers:2026 清单(1月–5月) — Sebastian Raschka
arXiv:2604.12374 LLM 基础设施 方法 OA · 绿色 被引 37 · S2

Nemotron 3 Super 是 Nemotron 3 系列中首个采用 NVFP4 进行预训练的模型,借助 LatentMoE(一种同时优化精度 per FLOP 与精度 per parameter 的新型 Mixture-of-Experts 架构),并集成 MTP 层以通过原生 speculative decoding 加速推理。Nemotron 3 Super is the first model in the Nemotron 3 family to be pre-trained in NVFP4, leverage LatentMoE, a new Mixture-of-Experts architecture that optimizes for both accuracy per FLOP and accuracy per parameter, and include MTP layers for inference acceleration through native speculative decoding.

CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing
CRISP:面向结构-质量驱动的悬崖感知输入自适应稀疏预填充
arXiv:2609.01925 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文用 C_struct 替代 Jensen-Shannon Divergence 路由:C_struct 是一种结构化代理,通过度量 Vertical-Slash 兼容位置上的 mass 来复现 JSD 的路由决策,同时消除池化 matmul 与后续 KL 散度开销。This work replaces the Jensen-Shannon Divergence routing with C_struct, a structural proxy that measures mass at Vertical-Slash compatible positions and reproduces JSD's routing decisions while eliminating both the pooled matmul and subsequent KL divergence overhead.

Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration
改变置信度,而非改变预测:面向后验校准的预测保持型修复
arXiv:2609.01072 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 CORD,第一个 post-fit adapter,通过从 calibrator 拟合中移除 preservation constraint 来修复完整校准概率向量,从而实现精确的预测保持,并将原始决策的精确恢复交由后续输出修复完成。CORD is proposed, the first post-fit adapter to impose exact prediction preservation by repairing the full calibrated probability vector by removing the preservation constraint from calibrator fitting and assigns exact recovery of the original decision to subsequent output repair.

Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance
优化中的渗流动力学:方差级联与离散尺度不变性
arXiv:2609.02373 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

通过将随机梯度流建模为渗流过程来研究随机梯度下降的动态,其中嵌套的架构对称性迫使子网络以离散块的形式合并,而非通过单边附着。The dynamics of Stochastic Gradient Descent is studied by modeling the stochastic gradient flow as a percolation process, in which nested architectural symmetries force subnetworks to merge in discrete blocks rather than by single-edge attachment.

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
锁于入口,开于内里:RLVR 收窄解空间的位置
arXiv:2608.29188 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

表面 prompting 未能恢复多样性,而针对入口的干预则成功:使用早期 checkpoint 的 late-layer parameter interpolation 在不损失 pass@1 的情况下将解的覆盖度提高了 37%。While surface prompting fails to recover diversity, entrance-targeted interventions succeed: late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1 and late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1.

4️⃣ Tangram · 多轮对话非均匀KV Cache — arXiv:2606.06302(⭐⭐⭐⭐ 新鲜 arXiv)
arXiv:2606.06302 LLM 基础设施 方法 OA · 绿色 被引 2 · S2

Tangram 是一种 serving 框架,将先前系统动态处理的内容静态解析,可作为现有 non-uniform 压缩方法的即插即用底座,在匹配其精度的同时,端到端吞吐量较 full-KV 基线最高提升 2.6×。Tangram is a serving framework that statically resolves what prior systems handle dynamically, and serves as a drop-in substrate for existing non-uniform compression methods, matching their accuracy while improving end-to-end throughput by up to $2.6\times over the full-KV baseline.

When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference
当量化破坏记忆时:低精度时序推理中的循环状态写回
arXiv:2609.04490 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

循环状态写回被确立为低精度循环动态的关键决定因素,state-storage interface 被识别为量化循环推理的核心设计考量。R recurrent-state write-back is established as a key determinant of low-precision recurrent dynamics and the state-storage interface is identified as a central design consideration for quantized recurrent inference.

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
BeaconKV:由 Beacon 查询引导的 KV Cache 压缩,用于高效大型推理模型推理
arXiv:2609.04971 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 BeaconKV,一种免训练的 KV cache 压缩方法,通过为每个全局查询簇维护紧凑代表性 beacon query 来预测哪些 KV 对将被重访,无需存储完整查询历史。BeaconKV is proposed, a training-free KV cache compression method that maintains beacon queries, compact representatives for each global query cluster, to anticipate which KV pairs will be revisited without storing the entire query history.

Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise
Encoded Early, Used Late:Transformer 在何处开始作用于所推断的对话伙伴专业度
arXiv:2609.07139 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

发现合作者专业能力在模型早期层最易被解码,到网络中点前降至接近随机水平,并基于合成语料以单个模型作为初步验证。It is found that partner expertise is most decodable in the early layers and falls to near chance before the midpoint of the network, and one model on a synthetic corpus is used as an initial demonstration.

Cadence: Error-Bounded Lossy Compression of Demand Time Series with a Time-Series Foundation Model
Cadence:基于时序基础模型的需求时序误差有界有损压缩
arXiv:2609.06008 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

我们提出 Cadence,一种面向数值时序的误差有界有损压缩器,将 3.3 亿参数的时序基础模型(Google TimesFM-3)与自适应算术编码器相结合,保证每个样本满足 |x_t - x̂_t| ≤ τ。一项负面结论限定了设计空间:在无损编码场景下,基础模型毫无价值,因为节省的比特数仅与预测器精度呈对数关系 Δb = log_2(MAE_old / MAE_new)。因此 TimesFM-3 相对 32 阶线性预测器 1.51 倍的精度优势,在 20.28 比特中仅换取 0.60 比特,中位数增益仅 +0.03%。误差有界编码仅在一点上突破了这一限制。We present Cadence, an error-bounded lossy compressor for numeric time series pairing a 330M-parameter time-series foundation model (Google TimesFM-3) with an adaptive arithmetic coder, guaranteeing |x_t-x_t|leτ on every sample. One negative result constrains the design space: for lossless coding a foundation model is worth nothing, because bits saved are logarithmic in predictor accuracy, Δb=log_2(MAE_{old}/MAE_{new}). So the 1.51times advantage TimesFM-3 holds over a 32-tap linear predictor buys 0.60 bits of 20.28, a median gain of +0.03%. Error-bounded coding escapes this at one point: once

3️⃣ Speculative Decoding 延迟可解释模型 — arXiv:2605.15051(⭐⭐⭐⭐ 调优必读)
arXiv:2605.15051 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文为 LLM serving 中的 SD(投机解码)提出了一种简单且可解释的 latency 模型,能够准确刻画实际观测到的 latency,解释为何加速比常随服务器负载上升而下降,并系统刻画了 draft length、acceptance rate 以及 verifier 与 drafter 规模在不同 serving 条件下对 latency 的影响。A simple and interpretable latency model for SD in LLM serving is developed that accurately describes observed latency, explains why speedups often diminish as server load increases, and characterizes how draft length, acceptance rate, and verifier-drafter size shape latency across serving conditions are characterized.

3. Tutti:让 SSD 后备 KV Cache 成为长上下文生产方案
arXiv:2605.03375 LLM 基础设施 方法 OA · 绿色 被引 6 · S2

Tutti 是一种高效的 SSD-backed KV caching 方案,将 CPU 从 HBM 与 SSD 之间的关键数据与 I/O 控制路径中彻底移除,在提供近乎无限容量的同时,实现了与 DRAM-backed LMCache 几乎相当的 inference 性能。Tutti is an efficient SSD-backed KV caching solution that eliminates CPU intervention from the critical data and I/O control paths between HBM and SSDs, and achieves nearly the same inference performance as DRAM-backed LMCache, while providing almost infinite capacity.

3. Flow-Controlled Scheduling for LLM Inference(arXiv 2604.11001)
Flow-Controlled Scheduling for LLM Inference(arXiv 2604.11001)
arXiv:2604.11001 LLM 基础设施 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出了一种简单的 flow-control 框架,通过控制 prompt 加入 LLM 活跃集合的速率,实现更高的 token 与 request 吞吐量、更低的平均与尾部 latency,以及更稳定的 KV cache 利用率。A simple flow-control framework is proposed that controls the rate at which prompts join the active set in large language models and achieves higher token and request throughput, lower average and tail latency, and more stable KV cache utilization.

Negative Self-Distillation: Learning to Reason by Avoiding Flaws
负自蒸馏:通过规避缺陷来学习推理
arXiv:2609.11699 LLM 基础设施 方法 OA · 绿色 被引 1 · S2

提出 Negative Self-Distillation(NSD),一个通过偏离有缺陷的推理而非模仿特权解来优化 LLM 的新框架,并一致优于 OPSD 及其他无标签、自举式强化学习(RL)基线。Negative Self-Distillation (NSD) is introduced, a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions, and consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.