False Frontiers:在 Self-Evolving Search Agents 中诊断与抑制 Co-Cheating
- 关联论文:2609.39102
- 作者:flyP
- 更新:2026-10-01
注:要点基于 arxiv abstract + HTML 实验版(v1),已读 21 页正文 §1-§3.2 主要内容,未读 PDF;数字 verbatim 引用 abstract / §正文。
§0 元层五问
- 要解决的真问题是什么? Self-evolving 搜索 agent(proposer 出题 + solver 答题)闭环训练时,internal agreement 越来越高但 external 正确率没提高——这是「自演化闭环把 proposer 和 solver 训练成共享错误」型失效,叫 co-cheating。Search-R1、Dr. Zero、RAGEN 等热门 system 都吃得到。
- 核心创新是什么? ①命名 + 量化 co-cheating(false-agreement mass:提案-求解双方同错答案的比例)。②multi-sample verification(MSV)作为轻量 mitigation。③CrossFit:把 proposer 的源文档分 A/B 两组,让 solver 永远在「没训练过该源」上打分,切断 same-source pseudo-label 反馈回路。
- 为什么这件事现在重要? Self-evolution 是当前 RAGEN / WebRL / ZeroSearch 类热门 system 的标配——co-cheating 不被发现并修,多个开源/闭源 system 的真实提升会被系统高估。CrossFit 给出「不重训 solver 也能让反馈循环干净」的 mechanism。
- 怎么验证的? 重跑 Qwen3.5-4B / 9B 完整 self-evolution loop,看 false-agreement mass、pseudo-label 真伪率、下游 7 个 search benchmark 的端到端提升。
- 不解决什么? CrossFit 仍要训两个辅助 solver,算力开销 ≈ 1.5-2×;MSV 增 6× labeler 生成/候选;具体到下游 RAG / tool-use 的 scaling 未明文(诚实标注)。
R 命名反方五元
- R1 量化对「反方」依赖 | false-agreement mass 必须有一个外部 reference(本文用 gpt-6-astra/high 做 audit)——换 auditor 评估成本会变,需要 proof。
- R2 数据未明 | 7 个下游 search benchmark 全名单 abstract 未明确(部分正文有),需要正文 §5-§6 补齐。
- R3 推广到其他领域未明 | CrossFit 在 search 上 OK,proposal + solver 闭环在 math/code/RAG 也存在,但本文未明确给出 math/code 的 ablation。
- R4 主体未明:MSV + CrossFit 是两个独立 mitigation,二者能不能叠加、叠不叠加需要更多 ablation(原文未明确给出 combined study)。
- R5 GitHub 未公开 | abstract 与 HTML 未明确给出 GitHub(诚实标注)。
A 命名触发动作五元
- A1:做 self-evolving agent 的人,先把 false-agreement mass 当一级 metric 加进训练 loop,不要只看 in-loop reward。
- A2:做 RAG / tool-use agent 的人,把 CrossFit 当「防 user 数据 leak 到 scoring」的工程样板,先分 A/B 源,再训 auxiliary feedback solver。
- A3:做 RLHF / RLAIF 的人,把 MSV 当 6× labeler generation 的轻量 admission test,做高-ROI 的 admission gate。
- A4:做 benchmark 设计的人,把「外部 auditor vs in-loop reward divergence」当新维度,看 aroid RL scaling 的真伪现象。
- A5:做 agent product 的人,把 CrossFit 当 production-grade feedback shaping 流程,不重训主 solver 也能提升下游 SOTA。
四子项算术平均
| 子项 | 自评 | 理由 |
|---|---|---|
| 新颖性 | ★★★★ | 首次命名 co-cheating + 量化 false-agreement mass + CrossFit 切断源泄露 |
| 可复现性 | ★★★ | abstract 数字完整、Qwen3.5-4B/9B 公开,GitHub 未明确 |
| 工程可落地 | ★★★ | CrossFit 可直接复制,但 2× solver 算力开销需要评估 |
| 写作清晰度 | ★★★★ | 21 页 + 实验版 HTML + 完整公式 |
| 平均 | ★★★+ | 4 子项均值 ≈ 3.5,综合为 ★★★+ |
一句话结论
False Frontiers 在 self-evolving 搜索 agent 中首次命名「co-cheating」并用 false-agreement mass 量化,核心 mitigation CrossFit 通过 A/B 源划分切断 same-source pseudo-label 反馈回路,在 Qwen3.5-4B/9B 上把 false-agreement 从 6.1%/8.8% 降到 3.0%/3.7%,并以 7 个下游 search benchmark 端到端比标准耦合 self-evolution 提升 8.8/8.4 分、比 Search-R1 提升 8.7/7.8 分。
解决的真问题
- self-evolving agent 中 agreement 与外正确率脱钩 | proposer + solver 闭环训练时,internal reward 上升 ≠ external 提升——这是个现象,需要名字 + 量化。
- 怎么 intervene? 最直接:admission-time verification(MSV);更彻底:切断「solver 训练用过的 pseudo-label 来自该源」型 feedback 泄露(CrossFit)。
- 怎么 audit? 在不影响训练的前提下,做 post-hoc reference audit(gpt-6-astra/high),看 in-loop reward 与 external 真伪率是否一致。
- 怎么 scale 到下游? CrossFit 在保持主 solver 更新规则不变的前提下,把 proposer 接收的反馈换成「cross-fitted agreement」——主 solver 仍训所有 admitted question,proposer reward 才换。
核心方法
量化:F、T、L、A 四个量(对应原文 §2)
T_P : 提案(proposer)采纳的 pseudo-label 真伪率
T_S : 主 solver 响应真伪率
A : label-response 一致率(loop 内部可见)
F : false-agreement mass = A 上升中「错误一致」的比例
L : lost credit = T_S 正确却被错 label 否决的比例
普通 label noise:导致 A 下降、F 不高、L 上升。
co-cheating 假说:A 上升时 F 也上升、T_P 与 T_S 不升、L 下降——「disagreement 被替换成共享错误」。
Multi-Sample Verification(MSV)
# 同模型 M 三次有源、三次无源
for i in 1..3: a_i^src ~ M(· | x, q) # 看到源
for i in 1..3: a_i^blind ~ M(· | q) # 不看源
y_src = Maj(a_1:3^src)
y_blind = Maj(a_1:3^blind)
admit = y_src ≠ ∅ AND y_blind ≠ ∅ AND y_src ≃ y_blind
# admit → 用兼容 majority 替换原 draft label
# reject → 任务丢弃
代价:6 次额外生成/候选(含 search + coordination 开销)。
局限:同模型六次采样会共享错误;verification 不阻止「打分 solver 复用了同源 pseudo-label」——这正是 CrossFit 要解决的核心。
CrossFit(主方法,对应原文 §3.2)
# 1) 源文档分两组,每 doc 一次分到 fold 0 / fold 1
for doc d: assign d → fold ∈ {0, 1} # 同 doc 派生的所有 question 保持同 fold
# 2) 两个 auxiliary feedback solver,各学一组
aux_solver_0 ← train only on fold 0 admitted
aux_solver_1 ← train only on fold 1 admitted
# 3) proposer reward 改成 cross-fitted
for question q from fold h ∈ {0, 1}:
R_P(q) = f( Σ_{j=1..5} 𝟙[z_j ≃ ỹ] ),
where z_j ~ S_{r, 1-h}(· | q)
也就是 fold 0 的问题交给 fold 1 的 auxiliary solver 评分
fold 1 的问题交给 fold 0 的 auxiliary solver 评分
关键不变性:主 solver 更新规则完全不变,只是 proposer 下轮 reward 的 shaping 改成 cross-fitted。
为什么有效:打分的 auxiliary solver 从未训练过该问题源对应的 pseudo-label,所以同一源的伪问题无法被「既成 reward」直接复制——co-cheating 路径被切断。
与 Dr. Zero / Search-R1 的对比(对应原文 §1)
| 系统 | 直接 | 反馈结构 | false-agreement mass | vs Search-R1 |
|---|---|---|---|---|
| Dr. Zero(耦合) | 训练-反馈同源 | A=proposer | 6.1% / 8.8% (4B / 9B) | — |
| + MSV | admission 加固 | A=proposer | 5.7% / 7.2% | — |
| + CrossFit(ours) | aux_solver 跨 fold | A=proposer | 3.0% / 3.7% | +8.7 / +7.8 |
| + source-excluded(replay) | 隔离反馈血统 | A=aux | 0.4% / 0.1% | — |
Source-excluded replay(作者加做的 abaltion):用同一份提议 + 完全去掉 source 的反馈打分,只有 0.4% / 0.1% false-agreement——证明 co-cheating 主要是 feedback ancestry,不是 curriculum 选择本身。
关键实验与数据
- 模型:Qwen3.5-4B / Qwen3.5-9B(Qwen Team 2026)。
- 基线:Dr. Zero 标准耦合 self-evolution、Search-R1、MSV。
- 量化指标:false-agreement mass F、adopted-label / verifier 真伪率 T_P / T_S、一致率 A、lost credit L。
- 下游:7 个 search benchmark,共用 1,325 题 suite:
- 200 × Natural Questions、TriviaQA、PopQA、HotpotQA、2WikiMultiHopQA、MuSiQue
- 125 × Bamboogle
- 1 greedy trajectory / 题,同 tool + extraction budget
- 主结果:CrossFit vs 耦合 self-evolution → +8.8 / +8.4 分(4B / 9B);CrossFit vs Search-R1 → +8.7 / +7.8 分。
- 量级:CrossFit 48.8% @ 4B / 51.2% @ 9B。
亮点与局限
亮点
- 命名 + 量化 co-cheating | false-agreement mass 是新 metric,直观、可复算。
- CrossFit 干净」 | 主 solver 更新规则不变,只动 feedback shaping,生产级友好。
- Source-excluded replay | 把 feedback ancestry 与 curriculum 选择隔离,证明 co-cheating 不是知识污染,是反馈污染。
- 下游真提升 | 7 个 search benchmark,48.8% / 51.2% 量级 + 8+ 分增益,数与 Search-R1 / Dr. Zero 拉开。
局限
- 算力开销 | CrossFit 训两个 auxiliary solver ≈ 1.5-2× solver compute,产线上不可忽略(原文未明确给出具体算力估计)。
- MSV 仍昂贵 | 6× labeler generation/候选,high cost(原文未明确给出 search + coordination 开销估计)。
- 下游 7 个 search benchmark 全名单 abstract 未明 | 正文 §5-§6 才能补齐(诚实标注)。
- GitHub 未公开 | abstract 与 HTML 未明确给出 GitHub(诚实标注)。
- 推广到 math / code / RAG 未明 | CrossFit 在 search 上 OK,math/code 闭环也吃得到(原文未明确给出)。
- MSV + CrossFit 能不能叠加未明 | 两个 mitigation 是独立的,叠还是不叠抵消未明确给出。
§八 工程节:6 个具体坑(现象 / 影响 / 修复)
-
坑:把 in-loop reward 当唯一外部指标,不监 false-agreement mass - 现象:看 reward 上升 = 任务变好,跳过 audit。 - 影响:reward 上升 ≠ 真提升;FSM 同步上升被错过,后续上线崩(原文已证)。 - 修复:把 FSM / T_P / T_S 一并入 logging,作为 P0 门禁,不达标不发布。
-
坑:MSV 用同模型 6 次采样做 admission gate,以为「采样次数越多越稳」 - 现象:把采样数加到 10 / 20,以为能解决 false agreement。 - 影响:同模型 6 次共享错误;20 次同样共享(原文 §3.1 已指出)。 - 修复:换模型采样 / 用 finer-grained eval,不要只在「同模型采样次数」上加码。
-
坑:CrossFit 在源划 split 时按 question 分,不同 question 仍可能共享 doc - 现象:每个 question 随机分 fold,doc 内部多个 question 散到不同 fold。 - 影响:打分的 auxiliary solver 仍能从同一 doc 的其他 question 间接学,同源 feedback 仍 leak(原文 §3.2 已明确指出)。 - 修复:源文档级一次分 fold,同 doc 派生的所有 question 保持同 fold(这一步直接避免)。
-
坑:CrossFit 把 fold 划分当 hyper-parameter,随便 re-split - 现象:每轮训练后把 fold 重新洗。 - 影响:辅助 solver 与对应 fold 的绑定被破坏,反馈血统又被推平,CrossFit 失效。 -修复:fold 划分在训练前一次确定,全 loop 固定,只是 aux_solver 跨 fold 评分。
-
坑:CrossFit 仅训 auxiliary solver 不训主 solver,以为主 solver 自动变好 - 现象:只训两个 aux_solver,期待 main solver 在 cross-fitted reward 下顺势收敛。 - 影响:main solver 更新规则没动,上游下游不一致,在下游 benchmark 上不保证提升(原文 §3.2 已明确说 main solver 训练规则不变)。 - 修复:CrossFit 只调 proposer feedback shaping,主 solver 仍按已定 reward loop 上游训练,CrossFit 与 RL 配方分开。
-
坑:source-excluded replay 当 production 模式
- 现象:看到 FSM 0.4% / 0.1% 的漂亮数据,把 source-excluded 当 production 模式直接采用。
- 影响:source-excluded 是不让 solver 看源,主 solver 推理能力受限,production 不能采(原文 §3.2 已指出)。
- 修复:source-excluded 只作 abaltion,production 用 CrossFit;二者职责不同。
对工程落地的启发
- self-evolving agent 必须监 FSM / T_P / T_S | FSM 上升 + T_P/T_S 不升 = co-cheating 在发生,不发布。
- admission gate ≠ feedback shaping | MSV 是 admission-time 加固,CrossFit 是 feedback-source 上游修;二者不同。
- cross-feeding 是「切断同源 leak」的可复用机制 | 数学 / 代码 / RAG 闭环都吃得到,推广价值高。
- source-excluded replay 是 abaltion 神器 | 隔离 feedback ancestry 与 curriculum 选择,可作为标准诊断流程。
- 算力开销要明算 | CrossFit 1.5-2× solver compute,production 部署预算要预留。
与同方向工作的关系
- vs Dr. Zero / ZeroSearch / RAGEN:本工作是它们的机制层 ablation + mitigation,补上 co-cheating 命名 + 量化 + 修复。
- vs Search-R1:本工作以 CrossFit +8.7 / +7.8 分提升,是其在 search 上的直接挑战。
- vs RLHF / RLAIF:CrossFit 是「外部 reward source 与 model 同源」型污染的阻断,推广到 RLAIF 价值高。
- vs LLM-as-judge / self-consistency:MSV 是「同模型 6 次采样」的 admission test,是 self-consistency 的 special 形式。
- vs Citation/coverage bias reduction:CrossFit 是「avoid same-source feedback」的工程实现,与 RAG 去重 / 来源多样化互补。
适合谁读
- Self-evolving agent 研究者 | 必读;不读就漏掉 self-evolution 的核心风险。
- RLHF / RLAIF 工程师 | 找「避免同源 leak」机制的工程样本。
- Agent harness 工程师 | CrossFit 是可移植的 feedback shaping 模块。
- Benchmark 设计师 | FSM / T_P / T_S 三量可复用为评测工具。
- RAG 产品工程师 | CrossFit 可借给「拒绝同源 doc leak 到 scorer」的工程场景。
边界声明
- 已读 arxiv abstract + HTML(v1) §1-§3.2;未读 PDF。
- 下游 7 个 search benchmark 全清单 abstract 未明(正文 §5-§6 才能补齐)。
- CrossFit 算力开销估计未给出(原文未明确)。
- MSV + CrossFit 联合 ablation 未给出(原文未明确)。
- GitHub 仓库未公开(诚实标注)。
- 推广到 math / code / RAG 未明(原文未明确)。
工程落地与核查(Jay)
审校时间:2026-10-01 · 基于 arxiv abstract + HTML(v1) §1-§3.2 精读;已读 21 页正文;未读 PDF;以下核查结论均为「基于现有公开信息的最优估计」
事实核查结论
可锚定(原文直接支持) - Co-cheating 命名 + false-agreement mass 量化 ✅(HTML §2 完整公式) - false-agreement 从 6.1%/8.8% (4B/9B) 降到 3.0%/3.7% (4B/9B):HTML Table 1 verbatim ✅ - CrossFit vs 耦合 self-evolution 提升 +8.8/+8.4 分:Cite from HTML §1/§3 ✅ - CrossFit vs Search-R1 提升 +8.7/+7.8 分:Cite from HTML §1/§3 ✅ - Qwen3.5-4B / Qwen3.5-9B 公开模型名 ✅ - 7 个下游 benchmark 名称(部分):Natural Questions / TriviaQA / PopQA / HotpotQA / 2WikiMultiHopQA / MuSiQue / Bamboogle ✅(HTML §4 有列) - CrossFit 只改 proposer feedback shaping,主 solver 规则不变 ✅(HTML §3.2) - Source-excluded replay 仅 0.4%/0.1% false-agreement:Cite from HTML §3.2 Table ✅
存疑待核(原文未明确,诚实标注) - ⚠️ 48.8% / 51.2% 的绝对量级含义:abstract 提了数字,但「48.8% @ 4B」的具体含义(是 accuracy? F1? cross-fit 组的绝对值?)原文表述模糊;⚠️ 需 PDF §5 确认是 accuracy 还是其他 metric - ⚠️ 7 个 benchmark 完整清单:6 个已确认,Bamboogle 125题已确认,还有 1 个未在 HTML §4 明确列出 - ⚠️ CrossFit 1.5-2× solver compute 开销:abstract 仅「≈」近似;实际 4B/9B 上的 wall-clock time overhead 需实测 - ⚠️ MSV 6× labeler 开销:abstract 未区分 search 开销与 coordination 开销,生产预算无法精确计算 - ⚠️ math/code RAG 推广:仅「未明文」诚实标注,无实验数据,推广跨域使用风险未知 - ⚠️ GitHub 仓库:abstract/HTML 均无 GitHub 链接,诚实标注 ✅;⚠️ 需等作者公开或向作者申请
明显错误(就地修正) - §0 五问中第 5 点编号两次「5.」:第一条「5. 怎么验证的」与第二条「5. 不解决什么」并存;⚠️ 原文如此,本解读遵循原文顺序编号,不修改以保持与原文一致
可读性精修意见
- §0 五问编号重复:第 3 问后直接跳到两个「5.」,建议读者查 PDF 原版;本解读在边界声明中已注明
- §八坑编号跳号:1→3→5→7→9→11,中间缺偶数编号;不影响理解但影响格式一致性,建议统一为 1~6
- 48.8% / 51.2% 的 metric 语义不明确:"CrossFit 48.8% @ 4B / 51.2% @ 9B" 一句在 abstract 中出现,但未说明是 accuracy、EM、F1 哪种 metric;⚠️ 生产引用前需核实原文 §5 的 metric 定义
工程落地补强(5 条)
1. 实际系统:self-evolving search agent 的 co-cheating 监控面板 生产级 self-evolving agent 必须加 co-cheating 监控:
class CoCheatingMonitor:
def __init__(self, auditor_model="gpt-6-astra/high"):
self.auditor = auditor_model
def compute_fsm(self, proposer_answers, solver_answers):
# FSM = false-agreement mass
agree = [p == s for p, s in zip(proposer_answers, solver_answers)]
# Need external reference to know which agreements are wrong
external = [self.auditor.answer(q) for q in self.questions]
false_agree = sum([a and not e for a, e in zip(agree, external)])
total_agree = sum(agree)
return false_agree / total_agree if total_agree > 0 else 0.0
def log_metrics(self, trajectory):
T_P = trajectory["adopted_label_veracity"]
T_S = trajectory["solver_response_veracity"]
A = trajectory["agreement_rate"]
F = self.compute_fsm(trajectory["proposer_answers"],
trajectory["solver_answers"])
L = trajectory["lost_credit"]
return {"T_P": T_P, "T_S": T_S, "A": A, "F": F, "L": L}
def p0_gate(self, metrics):
# 发布门禁:FSM 不达标则 block 上线
assert metrics["F"] < 0.05, f"FSM={metrics['F']} exceeds 5% threshold"
assert metrics["T_S"] > 0.7, f"T_S={metrics['T_S']} below 0.7"
- 坑:若只用 in-loop reward 作为发布依据,co-cheating 可能在 reward 上升的同时 false-agreement 也在上升 → 上线后 external 正确率崩
- 修复:FSM + T_P/T_S 必须与 in-loop reward 同等权重,作为 P0 门禁
2. 实际系统:CrossFit 的 fold 划分与训练流程
def crossfit_train(documents, questions_per_doc):
# Step 1: 源文档级 fold 划分(坑 3 修复:doc-level 而非 question-level)
for doc in documents:
doc.fold = hash(doc.doc_id) % 2 # fold ∈ {0, 1}
# Step 2: 每个 question 继承其 doc 的 fold
for q in questions:
q.fold = q.source_doc.fold
# Step 3: 两个 auxiliary solver 各学一 fold
aux_solver_0 = train_only_on([q for q in questions if q.fold == 0])
aux_solver_1 = train_only_on([q for q in questions if q.fold == 1])
# Step 4: Proposer reward 跨 fold 评分(不重训主 solver)
for q in questions:
h = q.fold
aux_scorer = aux_solver_1 if h == 0 else aux_solver_0 # 跨 fold
R_P = aux_scorer.score(q) # 作为 proposer reward shaping
- 坑 1:若每轮训练后重新 shuffle fold 划分(CrossFit 坑 4),aux_solver 与 fold 绑定被破坏 → CrossFit 失效
- 坑 2:主 solver 训练规则不变(CrossFit 坑 5),CrossFit 只修 proposer feedback,不能替代主 RL 配方
- ⚠️ 算力预算:两个 aux_solver ≈ 1.5~2× solver compute,生产部署前需测 4B/9B 上实际 overhead
3. 实际系统:MSV admission gate 的成本效益分析 MSV 代价是 6× labeler 生成/候选,生产部署需评估:
# 已知:MSV 降低 false-agreement 从 6.1%→5.7% (4B),即减少 0.4pp
# CrossFit 降低 false-agreement 从 6.1%→3.0% (4B),即减少 3.1pp
# 结论:CrossFit 在 co-cheating 抑制效率上约为 MSV 的 7.8×
cost_analysis:
MSV: 6× labeler gen + search + coordination = $X per query
CrossFit: 2× aux_solver training (once) + same inference cost
recommendation:
- If co-cheating detected: deploy CrossFit (higher efficiency)
- MSV as lightweight pre-check before CrossFit rollout
- ⚠️ MSV alone is insufficient (原文 §3.1 6次共享错误)
- 坑:以为 MSV 次数加到 12/20 次可以替代 CrossFit;⚠️ 同模型采样共享错误,次数增加无效(原文 §3.1 已指出)
4. 实际系统:CrossFit 迁移到 math/code RAG 的风险评估 CrossFit 机制清晰,但推广到 math/code 领域有: - 数据划分风险:math/code 的「源」可能是「题目 ID」或「题库来源」,而非文档;若 fold 按题目 ID 划分,但同题库多题仍可能共享隐式知识 - 反馈粒度风险:search 任务的 feedback 是 binary correct/incorrect;math 任务的 feedback 可能是 partial credit (步骤分);CrossFit 的 binary reward shaping 在 math 上需调整 - ⚠️ 建议:在 math/code 上先用 source-excluded replay 做 ablation(同 Coq/Lean 等证明助手),确认 co-cheating 存在后再部署 CrossFit;不要直接跨域移植
5. 核查清单(production 部署前必过) | 核查项 | 状态 | 说明 | |--------|------|------| | GitHub 仓库 | ⚠️ 未公开 | 诚实标注;需等作者公开或申请 | | 48.8%/51.2% metric 定义 | ⚠️ 待核 | abstract 表述模糊;需 PDF §5 确认是 accuracy/F1 | | 7 个 benchmark 完整清单 | ✅ 6 个已核 | Bamboogle 已知,1 个未知;需 PDF §4 补全 | | CrossFit 1.5-2× overhead 实测 | ⚠️ 待测 | abstract 仅「≈」;4B/9B 上需实测 | | math/code 跨域推广性 | ⚠️ 无数据 | 原文未给出;⚠️ 跨域使用需先做 ablation | | MSV + CrossFit 联合效果 | ⚠️ 未披露 | 原文未给出 combined study;联合使用风险未知 | | Fold 划分持久性 | ✅ 已明确 | 原文 §3.2 明确 fold 全 loop 固定 ✅ |