Scaling Properties of Same-Family On-Policy Distillation

  • 类型:arxiv
  • 标识:2609.32722
  • 链接:https://arxiv.org/abs/2609.32722
  • 主分类:engineering
  • 形态:method
  • TLDR:Reinforcement learning (RL) can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of on-policy distillation (OPD) across weak-to-strong, same-base, and strong-to-weak teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular useful-transfer regime, in which held-out accuracy (the gold score, G) rises approximately linearly in d=mathrm{KL(π_θVert π_{ref})}, the square root of token-level revers
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-01-agent-rag-longcontext-candidates.json