Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
- 类型:arxiv
- 标识:2608.29464
- 链接:https://arxiv.org/abs/2608.29464
- 主分类:evaluation
- 形态:method
- TLDR:Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness tests often place explicit bias cues in the user message, while agents may encounter preferences through tool returns or raw artifacts. We introduce FACE-Eval (Faithful Attribution of Cue Effects Evaluation), a 5,100-sample evaluation that varies cue location (user message or tool return) and explicitness (direct summary or raw artifact). We measure verbalized commitment among cue-following answers and unverbalized adoption among all cued samples
- 待LLM分类:否
- 标题中文:思维链忠实性随偏好线索的传递位置与方式而变化
- TLDR中文:思维链(CoT)监控假设推理轨迹能忠实记录影响模型答案的信息。现有忠实性测试通常将显式偏差线索放在用户消息中,而 Agent 可能通过工具返回值或原始产物遇到偏好。我们提出 FACE-Eval(线索效应忠实归因评估),一个 5,100 样本的评测方案,变化线索位置(用户消息或工具返回)与显隐程度(直接摘要或原始产物)。我们在跟随线索的答案中度量显式承诺度,并在所有含线索样本中度量未言明的采纳度
- 来源文件:
- /inbox/tom/_candidates/2026-09-01-agent-rag-longcontext-candidates.json