CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

  • 类型:arxiv
  • 标识:2607.25659
  • 链接:https://arxiv.org/abs/2607.25659
  • 主分类:evaluation
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar + OpenAlex
  • S2被引:0
  • OpenAlex被引:0
  • 影响力被引:0
  • TLDR:CoRT is proposed, a token-level credit weighting method for rubric-conditioned GRPO that uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt, and suggests that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation.
  • OpenAlex ID:W7171674157
  • OpenAlex DOI:10.48550/arxiv.2607.25659
  • DOI:10.48550/arxiv.2607.25659
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://doi.org/10.48550/arxiv.2607.25659
  • OpenAlex更新:2026-08-25
  • 待LLM分类:否
  • 标题中文:CoRT: 用于 token 级 rubric 引导策略优化的反事实回放
  • TLDR中文:本文提出 CoRT,一种用于 rubric 条件化 GRPO 的 token 级 credit 加权方法,通过反事实回放在原始 rubric 条件化 prompt 和匹配的免准则 prompt 下对同一样本响应重新打分,并表明策略内部反事实似然对比为响应内 credit 分配提供了有效的训练信号。
  • 来源文件
  • /inbox/tom/_candidates/2026-07-30-agent-rag-longcontext-candidates.json
  • [S2 enrich]
  • [OpenAlex backfill]