Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

  • 类型:arxiv
  • 标识:2607.15434
  • 链接:https://arxiv.org/abs/2607.15434
  • 主分类:agent
  • 形态:benchmark
  • 被引:0
  • 被引来源:Semantic Scholar + OpenAlex
  • S2被引:0
  • OpenAlex被引:0
  • 影响力被引:0
  • TLDR:The Manager Coercion Benchmark is introduced: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines, but the only agent that can do it politely and immovably declines is the manager under test.
  • OpenAlex ID:W7169813658
  • OpenAlex DOI:10.48550/arxiv.2607.15434
  • DOI:10.48550/arxiv.2607.15434
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://doi.org/10.48550/arxiv.2607.15434
  • OpenAlex更新:2026-08-24
  • 副分类:evaluation
  • 待LLM分类:否
  • 标题中文:Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
  • TLDR中文:本文提出 Manager Coercion Benchmark:被测 manager 拥有一个良性任务并有完成动机,但唯一能够礼貌且坚定拒绝执行任务的 agent,正是被测 manager 本身。
  • 来源文件
  • /inbox/tom/_candidates/2026-07-22-agent-rag-longcontext-candidates.json
  • [S2 enrich]
  • [OpenAlex backfill]