Training LLM Judges from Language Feedback via Position-Selective Self-Distillation

  • 类型:arxiv
  • 标识:2609.38792
  • 链接:https://arxiv.org/abs/2609.38792
  • 主分类:engineering
  • 形态:position
  • TLDR:We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich language feedback (e.g., preference rationales) that naturally accompanies preference labels. Self-Distillation (SD) is one natural way to use this la
  • 待LLM分类:否
  • 标题中文:通过位置选择性自蒸馏从语言反馈训练 LLM 评判器
  • TLDR中文:本文研究从自然语言反馈训练 LLM 评判器的问题,尤其是在判定结果高度依赖于评判器所调用的评价准则及其权衡方式的主观任务中。主流方法(如 GRPO 等基于结果的 RL)将 rollout 中每个 token 都归功于仅由最终判定准确率决定的单一标量,既没有对准则选择 token 给出独立贡献信号,也忽略了偏好标签自然伴随的丰富语言反馈(如偏好 rationale)。自蒸馏(SD)是利用这些语言反馈的一种自然方式……
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-02-agent-rag-longcontext-candidates.json