On-Policy Delta Distillation

  • 类型:arxiv
  • 标识:2607.15161
  • 链接:https://arxiv.org/abs/2607.15161
  • 主分类:multimodal
  • 形态:method
  • 被引:2
  • 被引来源:Semantic Scholar
  • S2被引:2
  • 影响力被引:0
  • TLDR:It is shown that the delta signal substantially improves on-policy distillation and the new distillation method is referred to as On-Policy Delta Distillation (OPD), enabling reasoning LLMs to achieve strong performance with only a short post-training period.
  • 待LLM分类:否
  • 标题中文:在策略差值蒸馏
  • TLDR中文:表明差值信号显著提升在策略蒸馏效果,提出新方法称为在策略差值蒸馏(OPD),使推理LLM仅需短暂的后训练即可获得强性能。
  • 来源文件
  • /inbox/tom/_candidates/2026-07-20-agent-rag-longcontext-candidates.json
  • [S2 enrich]