arXiv:2608.28389 · RAG 检索增强
CamoDocs: A Poisoning Attack Against Retrieval-Augmented Language Models Using Camouflaged Documents
CamoDocs:利用伪装文档针对检索增强语言模型的投毒攻击
CamoDocs: A Poisoning Attack Against Retrieval-Augmented Language Models Using Camouflaged Documents
- 类型:arxiv
- 标识:2608.28389
- 链接:http://arxiv.org/abs/2608.28389v1
- 主分类:rag
- 形态:method
- TLDR:Retrieval-augmented generation (RAG) augments LLMs with external documents, but public or user-editable sources expose RAG systems to data poisoning: attackers can inject malicious documents to steer outputs toward targeted answers. Existing poisoning attacks often rely on query inclusion, inserting the target query into poisoned documents to improve retrieval; however, this creates lexical and embedding-space artifacts that make them easy to filter. We propose CamoDocs, a poisoning attack that avoids direct query inclusion by camouflaging adversarial documents among benign content. CamoDocs c
- 副分类:risk
- 待LLM分类:否
- 标题中文:CamoDocs:利用伪装文档针对检索增强语言模型的投毒攻击
- TLDR中文:检索增强生成 (RAG) 通过外部文档增强 LLM,但公开或用户可编辑的数据源使 RAG 系统面临数据投毒风险:攻击者可注入恶意文档,将输出导向预设答案。现有投毒攻击多依赖查询包含 (query inclusion),将目标查询插入投毒文档以提升检索;但这会在词法与嵌入空间中留下痕迹,使其易于被过滤。本文提出 CamoDocs,一种通过在良性内容中伪装对抗性文档以避免直接包含查询的投毒攻击。CamoDocs c
- 来源文件:
- /inbox/tom/_candidates/2026-09-01-agent-rag-longcontext-candidates.json