WARP: A Word-Level Backdoor Attack Targeting RAG Systems via Retrieval Corpus Poisoning
Hui Liu, Yibo Zhou, Liguo Dong, Weidong Li, Shui Yu
2026Year
Abstract
Retrieval-Augmented Generation (RAG) systems retrieve relevant documents from a corpus database to mitigate issues like hallucination, outdated knowledge, and limited domain coverage. While enhancing large language models (LLMs) performance, RAG also introduces a new attack surface: adversaries can inject trigger-embedded malicious documents into the corpus database, potentially causing the LLM to produce attacker-controlled outputs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a259f4d9-de97-471b-b57d-b736f0adc7d0Related papers
- On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application DomainsXun Xian, Ganghua Wang, Xuan Bi, Rui Zhang et al.ICML 2025
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 6 citations
- Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker DocumentsAvital Shafran, Roei Schuster, Vitaly ShmatikovUSENIX Security 2025
- Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation SystemsHaowei Wang, Rupeng Zhang, Junjie Wang, Mingyang Li et al.AAAI 2026 · 3 citations
- AIP: Subverting Retrieval-Augmented Generation via Adversarial Instructional PromptSaket S. Chaturvedi, Gaurav Bagwe, Lan Zhang, Xiaoyong YuanEMNLP 2025
