JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

  • 类型:arxiv
  • 标识:2609.26550
  • 链接:https://arxiv.org/abs/2609.26550
  • 主分类:evaluation
  • 形态:method
  • TLDR:LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger gaps arise when judgments require checking
  • 待LLM分类:否
  • 标题中文:JEV-as-a-Judge:自信则接受,不确定则升级
  • TLDR中文:LLM-as-a-judge 支持跨任务评估,但在大规模场景下推理成本与置信度可靠性成为关键问题。本文研究仅做判断的 judge 能否提供经济高效的首轮判别,并识别何时需要更强的评估。在与十六种生成式与奖励模型 judge 的对比中(采用盲法人类裁定作为参照),我们发现 jev-as-a-judge 在普通偏好与有证据支撑的事实性任务上,与作为最强对照的 SOTA LLM judge 仅相差 3 个百分点,成本仅为后者的 0.36%。在需要核查
  • 来源文件
  • /inbox/tom/_candidates/2026-09-23-agent-rag-longcontext-candidates.json