When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
- 类型:arxiv
- 标识:2609.20511
- 链接:https://arxiv.org/abs/2609.20511
- 主分类:engineering
- 形态:method
- TLDR:We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify termination-token mismatch between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama, and Gemma, the two models can place their stopping probability on different EOS tokens, even when their declared stopping sets are identical. This mismatch can suppress the student's preferred termination action without reliably transferring the teacher-preferred alternative. We show that aligning t
- 待LLM分类:是
- 标题中文:当 EOS token 不一致时:理解 On-Policy 蒸馏中的长度膨胀
- TLDR中文:我们研究 on-policy 蒸馏 (OPD) 中的长度膨胀现象,即学生回答会变得过长,甚至耗尽生成预算。我们发现基础学生模型与训练后教师模型之间的终止 token 不匹配是该行为的重要来源。在 Qwen3、Llama 和 Gemma 上,两个模型可能将停止概率分配到不同的 EOS token 上,即便它们声明的停止集合相同。这种不匹配会抑制学生偏好的终止动作,同时无法可靠地传递教师偏好的替代动作。我们证明对齐 t...
- 来源文件:
- /inbox/tom/_candidates/2026-09-18-agent-rag-longcontext-candidates.json