Tracing Target Answers in Poisoned Retrieval Corpora via Token Influence Attribution

  • 类型:arxiv
  • 标识:2606.25721
  • 链接:http://arxiv.org/abs/2606.25721v1
  • 主分类:rag
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar + OpenAlex
  • S2被引:0
  • OpenAlex被引:0
  • 影响力被引:0
  • TLDR:TRACE is presented, a lightweight detection framework that identifies poisoning attacks by tracing answer-related tokens through token influence attribution, and first discovers recurrent high-influence keywords across retrieved documents and then performs a secondary verification to confirm their influence on model predictions.
  • OpenAlex ID:W7165921170
  • OpenAlex DOI:10.48550/arxiv.2606.25721
  • DOI:10.48550/arxiv.2606.25721
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://doi.org/10.48550/arxiv.2606.25721
  • OpenAlex更新:2026-07-19
  • 待LLM分类:否
  • 标题中文:通过 Token 影响归因追踪投毒检索语料中的目标答案
  • TLDR中文:本文提出 TRACE——一种通过 token 影响归因追踪答案相关 token 来识别投毒攻击的轻量检测框架;该方法首先发现跨检索文档的反复出现的高影响关键词,再经二次验证确认其对模型预测的影响。
  • 来源文件
  • /inbox/tom/_candidates/2026-06-25-agent-rag-longcontext-candidates.json
  • [S2 enrich]
  • [OpenAlex backfill]