Negative Self-Distillation: Learning to Reason by Avoiding Flaws

  • 类型:arxiv
  • 标识:2609.11699
  • 链接:https://arxiv.org/abs/2609.11699
  • 主分类:engineering
  • 形态:method
  • TLDR:On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self-corrective behaviors required to solve chal
  • 待LLM分类:是
  • 标题中文:负自蒸馏:通过规避缺陷来学习推理
  • TLDR中文:策略内自蒸馏(On-Policy Self-Distillation, OPSD)已成为大语言模型(LLM)自我改进的流行范式,允许模型通过利用真实解等特权信息充当自身的教师。然而近期研究表明,OPSD 可能严重损害 LLM 在复杂推理任务上的表现:通过强制学生模仿在特权信息条件下人为高置信度的推理轨迹,OPSD 在不经意间抑制了不确定性表达,并惩罚了解题所需的探索性与自我修正行为……
  • 来源文件
  • /inbox/tom/_candidates/2026-09-11-agent-rag-longcontext-candidates.json