Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

  • 类型:arxiv
  • 标识:2609.02998
  • 链接:https://arxiv.org/abs/2609.02998
  • 主分类:multimodal
  • 形态:method
  • TLDR:On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, without checking whether the teacher is reliable for each prompt. Because reverse KL is mode-seeking, a confidently wrong teacher can induce a strong yet misleading update. Distributional proxies, such as entropy or teacher-student likelihood agreement, measure uncertainty or agreement but do not directly verify outcome correctness. We introduce Teacher-Gated On-Policy Distillation (
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-08-agent-rag-longcontext-candidates.json