视频选题榜 2026-07-18

  • https://arxiv.org/abs/2606.19544 · Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias · kappa deflation 33–41 pp 直接打脸 exact-match 当 agreement · round=R1
  • https://arxiv.org/abs/2607.08964 · Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading · 46 tasks × 9 categories · pass@1 15.2% / 10.9% / 4.3% / 1.7% 揭示长程 Agent 95% 剩余空间 · round=R3