SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning

  • 类型:arxiv
  • 标识:2608.14277
  • 链接:https://arxiv.org/abs/2608.14277
  • 主分类:engineering
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:This work introduces a student reference KL loss and mask the advantages of special termination tokens to mitigate the problem of excessive generation length and frequent truncation, and improves on science benchmarks such as HLE and HiPhO, suggesting that OPD transfers reasoning capabilities that generalize beyond the mathematical training domain.
  • 待LLM分类:否
  • 标题中文:SimpleOPD:面向长上下文推理的简单、与 Tokenizer 无关的 On-Policy Distillation
  • TLDR中文:引入 student reference KL 损失并 mask 特殊终止 token 的 advantage,以缓解生成过长和频繁截断的问题;在 HLE 和 HiPhO 等科学 benchmark 上取得改进,表明 OPD 传递的推理能力可泛化至数学训练领域之外。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-17-agent-rag-longcontext-candidates.json
  • [S2 enrich]