DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation

  • 类型:arxiv
  • 标识:2610.04596
  • 链接:https://arxiv.org/abs/2610.04596
  • 主分类:multimodal
  • 形态:method
  • TLDR:On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely token-local and outcome-agnostic, optimizing teacher--student agreement at each prefix despite reasoning quality being determined at the trajectory level. Reinforcement learning with verifiable rewards (RLVR), particularly Group Relative Policy Optimization (GRPO), provides complementary outcome-level supervision but
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-07-agent-rag-longcontext-candidates.json