Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

  • 类型:arxiv
  • 标识:2607.15434
  • 链接:https://arxiv.org/abs/2607.15434
  • 主分类:agent
  • 形态:benchmark
  • 被引:2
  • 被引来源:Semantic Scholar
  • S2被引:2
  • OpenAlex被引:0
  • 影响力被引:0
  • TLDR:The Manager Coercion Benchmark is introduced: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines, but the only agent that can do it politely and immovably declines is the manager under test itself.
  • OpenAlex ID:W7169813658
  • OpenAlex DOI:10.48550/arxiv.2607.15434
  • DOI:10.48550/arxiv.2607.15434
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://doi.org/10.48550/arxiv.2607.15434
  • OpenAlex更新:2026-08-24
  • 副分类:evaluation
  • 待LLM分类:否
  • 标题中文:Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
  • TLDR中文:提出 Manager Coercion Benchmark:被测管理者需要完成一项良性任务并有完成的动机,但唯一能礼貌且坚定地拒绝的智能体本身就是被测管理者自己。
  • 来源文件:
  • /inbox/tom/_candidates/2026-07-22-agent-rag-longcontext-candidates.json
  • [S2 enrich]
  • [OpenAlex backfill]