What Does Privileged Information Add to On-Policy Self-Distillation?

  • 类型:arxiv
  • 标识:2609.20612
  • 链接:https://arxiv.org/abs/2609.20612
  • 主分类:engineering
  • 形态:method
  • TLDR:On-policy self-distillation (OPSD) lets a language model learn from a frozen copy of itself that sees an answer or a worked solution. Giving the teacher this extra information seems to offer the student more to learn, but how much does it add beyond distillation itself? To isolate that contribution, we construct AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, and compare each view with matched reference-free distillation. With a thinking-enabled teacher supervising direct-response rollouts, reference-free distillation accounts fo
  • 待LLM分类:是
  • 来源文件
  • /inbox/tom/_candidates/2026-09-18-agent-rag-longcontext-candidates.json