Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection

  • 类型:arxiv
  • 标识:2609.35932
  • 链接:https://arxiv.org/abs/2609.35932
  • 主分类:agent
  • 形态:method
  • TLDR:Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactly the same text, and because tokenization runs on the server, the defender rather than the attacker decides which one the model receives. We use this to measure how much of the injected instruction's authority comes from the reserved token's learned representation. Encoding the forged mark
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-09-30-agent-rag-longcontext-candidates.json