Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents

  • 类型:arxiv
  • 标识:2608.30322
  • 链接:https://arxiv.org/abs/2608.30322
  • 主分类:agent
  • 形态:benchmark
  • TLDR:Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely control whether an agent has access to those conventions. We introduce a knowledge-gated task-construction protocol that separates a task instruction from a compact artefact containing private conventions, reference tables, and utility operators. Construction-time provenance, byte-identical task instructions across the provided- and withheld-artefact conditions, leak audits, and executable witnesses make dependence on the artefact explicit and testable. Across fifteen calibration tas
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-03-agent-rag-longcontext-candidates.json