Safety of Latent Communication in Multi-Agent Systems

  • 类型:arxiv
  • 标识:2609.39788
  • 链接:https://arxiv.org/abs/2609.39788
  • 主分类:agent
  • 形态:method
  • TLDR:Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs
  • 副分类:risk
  • 待LLM分类:否
  • 标题中文:多 Agent 系统中潜在通信的安全
  • TLDR中文:潜在通信使多 Agent 系统直接在内部表征空间交换信息,从而降低基于文本通信的 token、计算与延迟开销。为此,研究者引入轻量级可训练链路,将发送方表征映射到接收方输入空间。本工作表明,即便链路训练本身无害,相比基于文本的通信也可能提升有害服从度,而底层的安全对齐 Agent 保持不变。攻击者可通过在有害 query-response 对上优化链路来放大此效应
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-01-agent-rag-longcontext-candidates.json