RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in Retrieval-Augmented Generation

  • 类型:arxiv
  • 标识:2608.13010
  • 链接:http://arxiv.org/abs/2608.13010v1
  • 主分类:rag
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:Joint deployment reduces attack success from 67.4% to 14.0% while retaining 41.3% F1 on unpoisoned retrieval, demonstrating practical protection at both corpus ingestion and query time without poison labels or trusted corpora.
  • 待LLM分类:否
  • 标题中文:RAGSieve:用于 RAG 中知识投毒检测的自参考局部对比
  • TLDR中文:联合部署可将攻击成功率从 67.4% 降至 14.0%,同时在未投毒检索上保留 41.3% 的 F1,证明无需投毒标签或可信语料库即可在语料入库和查询时提供实用的双重保护。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-14-agent-rag-longcontext-candidates.json
  • [S2 enrich]