Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

  • 类型:arxiv
  • 标识:2608.13760
  • 链接:https://arxiv.org/abs/2608.13760
  • 主分类:engineering
  • 形态:benchmark
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:It is found that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.
  • 待LLM分类:否
  • 标题中文:增强并不意味着可预测:思维模型中的推理行为
  • TLDR中文:发现面向推理的训练并未优先放大具有最高 Lift 的行为,这促使研究者采用过程级目标,以奖励经过校准且有依据的推理,而不仅仅是表面形式。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-18-rag-retrieval-reranking-candidates.json
  • /inbox/tom/_candidates/2026-08-18-agent-rag-longcontext-candidates.json
  • [S2 enrich]