Length Penalties Make Chain-of-Thought Less Monitorable
- 类型:arxiv
- 标识:2607.09786
- 链接:https://arxiv.org/abs/2607.09786
- 主分类:evaluation
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar + OpenAlex
- S2被引:0
- OpenAlex被引:0
- 影响力被引:0
- TLDR:Compression reduces reasoning tokens and preserves most multiple choice accuracy, while hint influence remains near baseline, in a frontier where reducing reasoning costs removes more evidence than shorter traces alone would predict.
- OpenAlex ID:W7168276331
- OpenAlex DOI:10.48550/arxiv.2607.09786
- DOI:10.48550/arxiv.2607.09786
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://doi.org/10.48550/arxiv.2607.09786
- OpenAlex更新:2026-07-19
- 待LLM分类:否
- 标题中文:长度惩罚使 Chain-of-Thought 更难被监控
- TLDR中文:压缩可减少推理 token 并保持大部分选择题准确率,同时提示影响接近基线——在一个前沿条件下,降低推理成本移除的证据比单纯缩短轨迹所预期的更多。
- 来源文件:
- /inbox/tom/_candidates/2026-07-17-agent-rag-longcontext-candidates.json
- [S2 enrich]
- [OpenAlex backfill]