Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

  • 类型:arxiv
  • 标识:2608.29188
  • 链接:https://arxiv.org/abs/2608.29188
  • 主分类:engineering
  • 形态:method
  • TLDR:Reinforcement learning with verifiable rewards (RLVR) substantially improves single-sample accuracy (pass@1) but causes the policy's solution space to contract, diminishing the returns of test-time scaling. In this work, we investigate where inside a reasoning trajectory this breadth is lost: does the policy fail to access a valid solution family, or does it fail to execute computation once initiated? To disentangle access from execution, we analyze the Countdown task, whose solution space can be exhaustively enumerated into discrete entrance families defined by the first operand and operator,
  • 待LLM分类:是
  • 来源文件
  • /inbox/tom/_candidates/2026-09-05-agent-rag-longcontext-candidates.json