Length Penalties Make Chain-of-Thought Less Monitorable
- 类型:arxiv
- 标识:2607.09786
- 链接:https://arxiv.org/abs/2607.09786
- 主分类:evaluation
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar + OpenAlex
- S2被引:0
- OpenAlex被引:0
- 影响力被引:0
- TLDR:This work trains reasoning models with length penalties to curb overthinking and cut inference cost but shows that these penalties make the chain of thought less monitorable, and trades monitorability for inference cost by removing the evidence monitors depend on.
- OpenAlex ID:W7168276331
- OpenAlex DOI:10.48550/arxiv.2607.09786
- DOI:10.48550/arxiv.2607.09786
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://doi.org/10.48550/arxiv.2607.09786
- OpenAlex更新:2026-07-19
- 待LLM分类:否
- 标题中文:长度惩罚使 Chain-of-Thought 更难被监控
- TLDR中文:本工作通过对推理模型施加长度惩罚以抑制过度思考并降低推理成本,但表明这些惩罚会削弱思维链的可监控性,通过移除监控所依赖的证据,以可监控性换取推理成本。
- 来源文件:
- /inbox/tom/_candidates/2026-07-17-agent-rag-longcontext-candidates.json
- [S2 enrich]
- [OpenAlex backfill]