Selecting Diverse SFT Traces Improves Post-RL Generalization

  • 类型:arxiv
  • 标识:2609.33780
  • 链接:https://arxiv.org/abs/2609.33780
  • 主分类:engineering
  • 形态:method
  • TLDR:Verified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL). We present a comprehensive study of route diversity, the variation in the sequences of reasoning steps in supervised fine-tuning (SFT) data, and propose a lightweight, rule-based fingerprint to select for it. From one pool at one budget, with matched training recipes and checkpoints, selecting diverse rather than similar routes improves post-RL problem coverage across puzzles and mathematics, including on problems harder than those seen in either training stage. In synthetic experiments, r
  • 待LLM分类:否
  • 标题中文:选择多样化的 SFT 轨迹可提升 RL 后的泛化能力
  • TLDR中文:经验证的解题数据在为推理模型准备强化学习(RL)方面并非同等有效。我们系统性地研究了路线多样性,即监督微调(SFT)数据中推理步骤序列的差异程度,并提出一种轻量级的基于规则的指纹方法对其进行筛选。在同一解答池与相同预算下,使用一致的训练方案与检查点,选择多样化而非相似的路线,可提升 RL 之后在谜题与数学任务上的解题覆盖率,包括在难度超过任一训练阶段所见题目的问题上。在合成实验中,r
  • 来源文件:
  • /inbox/tom/_candidates/2026-09-30-agent-rag-longcontext-candidates.json