DataFlex-RL: An Evaluation Platform for RLVR Data Policies

  • 类型:arxiv
  • 标识:2609.06107
  • 链接:https://arxiv.org/abs/2609.06107
  • 主分类:evaluation
  • 形态:benchmark
  • TLDR:Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. We introduce DataFlex-RL, an evaluation platform for comparing these choices under a common GRPO recipe. Our primary experiment evaluates 13 configurations across 12 matched seeds using Qwen2.5-7B-Base and 12 mathematics, logic, and science benchmarks. Uniform GRPO improves the domain-balanced average accuracy by 7.76 percentage points over the untrained checkpoint. None of the eight rollout-selec
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-14-agent-rag-longcontext-candidates.json