What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals

  • 类型:arxiv
  • 标识:2608.19269
  • 链接:https://arxiv.org/abs/2608.19269
  • 主分类:llm-infra
  • 形态:position
  • TLDR:Evaluation artifacts specify a forward computation: a task, scorer, and reported metric. They do not necessarily license the claim attached to that metric because the historical evidence and alternative semantics needed to replay it may be unbound. We formalize this missing claim-replay layer through a frozen substrate D, a grounded family F, a claim query q, and the resulting identified set. We then census all 124 mechanically eligible Inspect Evals units at a pinned commit. Every unit receives a terminal disposition; 110 stop before deterministic inference because required historical evidenc
  • 副分类:evaluation
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-08-29-agent-rag-longcontext-candidates.json