One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

  • 类型:arxiv
  • 标识:2608.25936
  • 链接:https://arxiv.org/abs/2608.25936
  • 主分类:multimodal
  • 形态:survey
  • TLDR:On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Distillation (OPSD) removes that cost. The teacher is the model itself, conditioned on privileged information the student will not have at test time, such as a reference solution, a plan, or environment feedback. The teacher is no stronger than the student, only better informed. Early results were pro
  • 待LLM分类:否
  • 标题中文:一个症状,三个杠杆:On-Policy Self-Distillation 的批判性综述
  • TLDR中文:On-policy distillation 在教师对模型自身生成进行逐 token 打分的同时训练语言模型,兼具模仿学习的密集监督与强化学习的 on-policy 采样,但需要一个更大的模型担任教师。On-Policy Self-Distillation(OPSD)消除了这一开销:教师即模型本身,并以学生测试时无法获取的特权信息(如参考答案、计划或环境反馈)为条件。教师并不比学生更强,只是信息更充分。早期结果颇具前景……
  • 来源文件
  • /inbox/tom/_candidates/2026-09-08-agent-rag-longcontext-candidates.json