DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation
- 类型:arxiv
- 标识:2610.04596
- 链接:https://arxiv.org/abs/2610.04596
- 主分类:multimodal
- 形态:method
- TLDR:On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely token-local and outcome-agnostic, optimizing teacher--student agreement at each prefix despite reasoning quality being determined at the trajectory level. Reinforcement learning with verifiable rewards (RLVR), particularly Group Relative Policy Optimization (GRPO), provides complementary outcome-level supervision but
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-10-07-agent-rag-longcontext-candidates.json