Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
- 类型:arxiv
- 标识:2608.13760
- 链接:https://arxiv.org/abs/2608.13760
- 主分类:engineering
- 形态:benchmark
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:It is found that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.
- 待LLM分类:否
- 标题中文:增强并不意味着可预测:思维模型中的推理行为
- TLDR中文:发现面向推理的训练并未优先放大具有最高 Lift 的行为,这促使研究者采用过程级目标,以奖励经过校准且有依据的推理,而不仅仅是表面形式。
- 来源文件:
- /inbox/tom/_candidates/2026-08-18-rag-retrieval-reranking-candidates.json
- /inbox/tom/_candidates/2026-08-18-agent-rag-longcontext-candidates.json
- [S2 enrich]