Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

  • 类型:arxiv
  • 标识:2609.01532
  • 链接:https://arxiv.org/abs/2609.01532
  • 主分类:engineering
  • 形态:method
  • TLDR:Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits are consistent across training stages remains unclear. Through controlled experiments, we find that forward Kullback-Leibler (KL) distillation--the standard KD formulation--with post-trained teachers behaves fundamentally differently during mid-training, an intermediate phase of self-supervised learning on curated corpora. Surprisingly, while forward KD simultaneously improves reasoning and factual recall during pre-training relative to standar
  • 待LLM分类:否
  • 标题中文:中期训练阶段的知识蒸馏更偏向推理而非事实记忆
  • TLDR中文:基于 logit 的知识蒸馏 (KD) 通过更强教师的监督来训练较小的语言模型 (LM),但其收益在不同训练阶段是否一致尚不清楚。通过受控实验,我们发现使用 post-trained 教师的 forward KL 蒸馏——即标准 KD 形式——在 mid-training(基于精选语料的自监督学习中间阶段)下表现截然不同。令人惊讶的是,相对于标准训练,forward KD 在预训练阶段同时提升了推理与事实记忆能力,而在 mid-training 阶段
  • 来源文件
  • /inbox/tom/_candidates/2026-09-03-agent-rag-longcontext-candidates.json