Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability
- 类型:arxiv
- 标识:2610.03509
- 链接:https://arxiv.org/abs/2610.03509
- 主分类:engineering
- 形态:method
- TLDR:Chain-of-thought (CoT) reasoning allows humans to inspect how large language models reach their answers, and oversee model behaviour. This reasoning comes at an increased inference cost, motivating efficient methods that train models to solve tasks using fewer tokens. However, a common concern is that such training may cause models to skip important reasoning steps, so the CoT no longer faithfully reflects the model's decision. It is unclear whether or when this occurs in practice, since different efficiency methods apply length pressure to models' CoT in distinct ways, and faithfully explaini
- 待LLM分类:否
- 标题中文:高效推理训练并不总是损害 CoT 忠实性与可监控性
- TLDR中文:思维链(CoT)推理使人类能够检查大语言模型如何得出答案,并监督其行为。这种推理带来更高的推理成本,促使人们开发训练模型以更少 token 解题的高效方法。然而,一个常见的担忧是此类训练可能导致模型跳过关键推理步骤,使 CoT 不再忠实反映模型的决策。该现象在实践中是否发生、何时发生尚不明确,因为不同的高效方法以不同方式对模型 CoT 施加长度压力,且忠实
- 来源文件:
- /inbox/tom/_candidates/2026-10-05-agent-rag-longcontext-candidates.json