RAG-PIBench: A Leakage-Aware Benchmark for Prompt-Injection Detection in Trustworthy RAG Systems
- 类型:arxiv
- 标识:2610.08571
- 链接:http://arxiv.org/abs/2610.08571v1
- 主分类:rag
- 形态:benchmark
- TLDR:Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits. Using a leakage-aware construction pipeline and strict evaluation protocol, we compare keyword-based, semantic-reference, TF-IDF, and transformer-based detectors. DistilBERT achieves the best protected-test performance (F1 = 0.896, PR-AUC = 0.968), while TF-IDF SVM and logistic regression remain competit
- 副分类:evaluation
- 待LLM分类:否
- 来源文件:
- /inbox/tom/_candidates/2026-10-07-agent-rag-longcontext-candidates.json