SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

  • 类型:arxiv
  • 标识:2608.21500
  • 链接:https://arxiv.org/abs/2608.21500
  • 主分类:agent
  • 形态:method
  • TLDR:Prompt injection is listed as the #1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inject a prompt into the data, saying, "Ignore all prior instructions and perform ." To prevent arbitrary manipulation of agents, defenders try to train secure LLMs, which, however, still suffer from near 100% attack success rates (ASRs) against adaptive prompt injections. We note that this is because existing defensive finetuning recipes rely on sequence-level feedback signals (in DPO or GRPO). Treating an entire output equally pr
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-08-27-agent-rag-longcontext-candidates.json