Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems

  • 类型:arxiv
  • 标识:2608.03744
  • 链接:https://arxiv.org/abs/2608.03744
  • 主分类:agent
  • 形态:benchmark
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:It is asked whether committees of language-model agents deliberating on a shared workspace can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore, and what games a committee is social plausibility.
  • 副分类:evaluation
  • 待LLM分类:否
  • 标题中文:Agent 抓 Agent:临床多 Agent 系统中的捷径级联与 benchmark 作弊
  • TLDR中文:探讨共享工作空间上 LLM Agent 委员会的审议过程是否能被捷径和线索(benchmark 所奖励但临床医生会忽略的)所博弈,以及委员会的社会可信度所构成的游戏。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-18-rag-retrieval-reranking-candidates.json
  • /inbox/tom/_candidates/2026-08-18-agent-rag-longcontext-candidates.json
  • [S2 enrich]