Tracing Target Answers in Poisoned Retrieval Corpora via Token Influence Attribution
- 类型:arxiv
- 标识:2606.25721
- 链接:http://arxiv.org/abs/2606.25721v1
- 主分类:rag
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar + OpenAlex
- S2被引:0
- OpenAlex被引:0
- 影响力被引:0
- TLDR:TRACE is presented, a lightweight detection framework that identifies poisoning attacks by tracing answer-related tokens through token influence attribution, and first discovers recurrent high-influence keywords across retrieved documents and then performs a secondary verification to confirm their influence on model predictions.
- OpenAlex ID:W7165921170
- OpenAlex DOI:10.48550/arxiv.2606.25721
- DOI:10.48550/arxiv.2606.25721
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://doi.org/10.48550/arxiv.2606.25721
- OpenAlex更新:2026-07-19
- 待LLM分类:否
- 标题中文:通过 Token 影响归因追踪投毒检索语料中的目标答案
- TLDR中文:本文提出 TRACE——一种通过 token 影响归因追踪答案相关 token 来识别投毒攻击的轻量检测框架;该方法首先发现跨检索文档的反复出现的高影响关键词,再经二次验证确认其对模型预测的影响。
- 来源文件:
- /inbox/tom/_candidates/2026-06-25-agent-rag-longcontext-candidates.json
- [S2 enrich]
- [OpenAlex backfill]