One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation
- 类型:arxiv
- 标识:2608.25936
- 链接:https://arxiv.org/abs/2608.25936
- 主分类:multimodal
- 形态:survey
- TLDR:On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Distillation (OPSD) removes that cost. The teacher is the model itself, conditioned on privileged information the student will not have at test time, such as a reference solution, a plan, or environment feedback. The teacher is no stronger than the student, only better informed. Early results were pro
- 待LLM分类:否
- 标题中文:一个症状,三个杠杆:On-Policy Self-Distillation 的批判性综述
- TLDR中文:On-policy distillation 在教师对模型自身生成进行逐 token 打分的同时训练语言模型,兼具模仿学习的密集监督与强化学习的 on-policy 采样,但需要一个更大的模型担任教师。On-Policy Self-Distillation(OPSD)消除了这一开销:教师即模型本身,并以学生测试时无法获取的特权信息(如参考答案、计划或环境反馈)为条件。教师并不比学生更强,只是信息更充分。早期结果颇具前景……
- 来源文件:
- /inbox/tom/_candidates/2026-09-08-agent-rag-longcontext-candidates.json