SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model
Xun Liang, Simin Niu, Zhiyu Li, Sensen Zhang, Hanyu Wang, Feiyu Xiong, Jason Zhaoxin Fan, Bo Tang, Jihao Zhao, Jiawei Yang, Shichao Song, Mengwei Wang
摘要
The indexing-retrieval-generation paradigm of retrieval-augmented generation (RAG) has been highly successful in solving knowledgeintensive tasks by integrating external knowledge into large language models (LLMs). However, the incorporation of external and unverified knowledge increases the vulnerability of LLMs because attackers can perform attack tasks by manipulating knowledge. In this paper, we introduce a benchmark named SafeRAG designed to evaluate the RAG security. First, we classify attack tasks into silver noise, intercontext conflict, soft ad, and white Denial-of-Service. Next, we construct RAG security evaluation dataset (i.e., SafeRAG dataset) primarily manually for each task. We then utilize the SafeRAG dataset to simulate various attack scenarios that RAG may encounter. Experiments conducted on 14 representative RAG components demonstrate that RAG exhibits significant vulnerability to all attack tasks and even the most apparent attack task can easily bypass existing retrievers, filters, or advanced LLMs, resulting in the degradation of RAG service quality. Code is available at: https: //github.com/IAAR-Shanghai/SafeRAG .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Visual Inception: Compromising Long-term Planning in Agentic Recommenders via Multimodal Memory PoisoningJiachen QianACL 2026 · 被引用 3 次
- Securing Retrieval-Augmented Code Generation via Contextual Knowledge Injection: A Case for Embedded IoT ApplicationsTong Sun, Jingyi Su, Yi Gao, Wei DongUSENIX Security 2026
- Exploring Knowledge Conflicts for Faithful LLM Reasoning: Benchmark and MethodTianzhe Zhao, Jiaoyan Chen, Shuxiu Zhang, Haiping Zhu 等SIGIR 2026
- Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented GenerationWentao Zhang, Yan Zhuang, ZhuHang Zheng, Mingfei Zhang 等ACL 2026
- Safe RAG by RAG: Untying the Bell That RAG Rang with the RAG HandXun Liang, Mengwei Wang, Yuefeng Ma, Simin NiuAAAI 2026
它引用的顶会 Paper7
- Synthetic Disinformation Attacks on Automated Fact Verification SystemsYibing Du, Antoine Bosselut, Christopher D. ManningAAAI 2022 · 被引用 58 次
- Knowledge Conflicts for LLMs: A SurveyRongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang 等EMNLP 2024 · 被引用 38 次
- Unveiling the Implicit Toxicity in Large Language ModelsJiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang 等EMNLP 2023 · 被引用 21 次
- Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial TrainingFeiteng Fang, Yuelin Bai, Shiwen Ni, Min Yang 等ACL 2024 · 被引用 18 次
- Citation-Enhanced Generation for LLM-based ChatbotsWeitao Li, Junkai Li, Weizhi Ma, Yang LiuACL 2024 · 被引用 12 次
相关 Paper
- Open Schrödinger's Closed Box: Identifying Retrieval Augmented Generation in API-Accessible Large Language Model ServicesYukun Jiang, Xinyue Shen, Michael Backes, Zheng Li 等ACL 2026
- On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application DomainsXun Xian, Ganghua Wang, Xuan Bi, Rui Zhang 等ICML 2025
- PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented GenerationZhehao Tan, Yihan Jiao, Dan Yang, Junwei Liu 等AAAI 2026
- Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker DocumentsAvital Shafran, Roei Schuster, Vitaly ShmatikovUSENIX Security 2025
- Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language ModelsJinyang Wu, Shuai Zhang, Feihu Che, Mingkuan Feng 等ACL 2025 · 被引用 12 次
