Learning from Teacher Continuations at Student States
- 类型:arxiv
- 标识:2609.36246
- 链接:https://arxiv.org/abs/2609.36246
- 主分类:multimodal
- 形态:method
- TLDR:We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillation methods: (1) sequential covariate shift in offline supervised fine-tuning (SFT) on fixed teacher trajectories, (2) fragmented supervision under prefix failure in token-level on-policy distillation (OPD), and (3) the need for access to teacher token probabilities in distribution
- 待LLM分类:否
- 标题中文:在学生状态处从教师续写中学习
- TLDR中文:我们提出 OLIVE(OnLine InterVEntion)。在每次迭代中,演进中的学生策略生成新前缀,由教师自回归地续写,随后学生根据教师生成 token 上的交叉熵进行更新。每一项设计选择都针对性地解决现有蒸馏方法的相应局限:(1) 离线监督微调(SFT)在固定教师轨迹上的序列协变量偏移;(2) token 级在线策略蒸馏(OPD)中前缀失败导致的监督碎片化;(3) 分布匹配方法对教师 token 概率的依赖。
- 来源文件:
- /inbox/tom/_candidates/2026-09-30-agent-rag-longcontext-candidates.json