HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

  • 类型:arxiv
  • 标识:2607.25398
  • 链接:http://arxiv.org/abs/2607.25398v1
  • 主分类:agent
  • 形态:benchmark
  • 被引:0
  • 被引来源:Semantic Scholar + OpenAlex
  • S2被引:0
  • OpenAlex被引:0
  • 影响力被引:0
  • TLDR:HANDBOOK_md is presented, a benchmark of 65 agentic tasks modeled on how employees follow company handbooks, and every task modifies one of 10 base handbooks, altering the specific rules and thresholds on which grading depends, so no two tasks share the same set of policies.
  • OpenAlex ID:W7171644820
  • OpenAlex DOI:10.48550/arxiv.2607.25398
  • DOI:10.48550/arxiv.2607.25398
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://doi.org/10.48550/arxiv.2607.25398
  • OpenAlex更新:2026-08-25
  • 副分类:evaluation
  • 待LLM分类:否
  • 标题中文:HANDBOOK.md:面向长上下文 Agent 指令遵循的基准
  • TLDR中文:介绍 HANDBOOK_md,一个包含 65 个 agentic 任务的 benchmark,模拟员工遵循公司手册的方式;每项任务对 10 份基础手册之一进行修改,变动评分所依赖的具体规则和阈值,因此没有任何两个任务共享同一套策略。
  • 来源文件
  • /inbox/tom/_candidates/2026-07-30-agent-rag-longcontext-candidates.json
  • [S2 enrich]
  • [OpenAlex backfill]