Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
- 类型:arxiv
- 标识:2608.03744
- 链接:https://arxiv.org/abs/2608.03744
- 主分类:agent
- 形态:benchmark
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:It is asked whether committees of language-model agents deliberating on a shared workspace can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore, and what games a committee is social plausibility.
- 副分类:evaluation
- 待LLM分类:否
- 标题中文:Agent 抓 Agent:临床多 Agent 系统中的捷径级联与 benchmark 作弊
- TLDR中文:探讨共享工作空间上 LLM Agent 委员会的审议过程是否能被捷径和线索(benchmark 所奖励但临床医生会忽略的)所博弈,以及委员会的社会可信度所构成的游戏。
- 来源文件:
- /inbox/tom/_candidates/2026-08-18-rag-retrieval-reranking-candidates.json
- /inbox/tom/_candidates/2026-08-18-agent-rag-longcontext-candidates.json
- [S2 enrich]