Learning from Teacher Continuations at Student States

  • 类型:arxiv
  • 标识:2609.36246
  • 链接:https://arxiv.org/abs/2609.36246
  • 主分类:multimodal
  • 形态:method
  • TLDR:We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillation methods: (1) sequential covariate shift in offline supervised fine-tuning (SFT) on fixed teacher trajectories, (2) fragmented supervision under prefix failure in token-level on-policy distillation (OPD), and (3) the need for access to teacher token probabilities in distribution
  • 待LLM分类:否
  • 标题中文:在学生状态处从教师续写中学习
  • TLDR中文:我们提出 OLIVE(OnLine InterVEntion)。在每次迭代中,演进中的学生策略生成新前缀,由教师自回归地续写,随后学生根据教师生成 token 上的交叉熵进行更新。每一项设计选择都针对性地解决现有蒸馏方法的相应局限:(1) 离线监督微调(SFT)在固定教师轨迹上的序列协变量偏移;(2) token 级在线策略蒸馏(OPD)中前缀失败导致的监督碎片化;(3) 分布匹配方法对教师 token 概率的依赖。
  • 来源文件:
  • /inbox/tom/_candidates/2026-09-30-agent-rag-longcontext-candidates.json