MInTRL: Off-policy Intervention can boost On-policy RL

  • 类型:arxiv
  • 标识:2609.12419
  • 链接:https://arxiv.org/abs/2609.12419
  • 主分类:engineering
  • 形态:method
  • TLDR:Reinforcement learning with verifiable rewards is typically performed on-policy, keeping training data close to the current policy but limiting learning to trajectories that the policy can discover itself. Off-policy methods such as supervised fine-tuning, on the other hand, can leverage external knowledge beyond the base model's capabilities, but may suffer from large distribution shift. The key challenge is thus to expand exploration without sacrificing learnability. In this work, we introduce Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier throug
  • 待LLM分类:否
  • 标题中文:MInTRL:Off-policy 干预可增强 On-policy RL
  • TLDR中文:带可验证奖励的强化学习通常以 on-policy 方式进行,训练数据保持接近当前策略,但学习仅限于策略自身可发现的轨迹。Off-policy 方法如监督微调则可利用基础模型能力之外的外部知识,但可能存在较大的分布偏移。核心挑战在于扩展探索的同时不牺牲可学习性。本工作提出 Minimal Intervention Reinforcement Learning (MInTRL),通过 throug
  • 来源文件
  • /inbox/tom/_candidates/2026-09-15-agent-rag-longcontext-candidates.json