SilentRetrieval: Hijacking Retrieval-Augmented Generation via Semantically-Preserving Adversarial Data Poisoning
Jiachen Qian
Abstract
Retrieval-Augmented Generation (RAG) mitigates LLM hallucinations but introduces a critical vulnerability: corpus integrity. We present SilentRetrieval, a two-stage data poisoning attack that hijacks RAG systems through adversarially crafted yet fluent documents. Stage 1 introduces Coordinated Beam Search (CBS), a multi-token joint optimization with a penalized fluency-similarity objective that preconditions a topically relevant host document to remain retrievable after payload insertion while constraining perplexity. Stage 2 employs Context-Adaptive Trigger Generation (CATG), a lightweight trigger-fusion step that uses a frozen LLM to generate triggers contextually integrated with document content. Under a one-poisoned-document-per-query evaluation with synthetic target answers, SilentRetrieval achieves 84.6%/81.3% HR@10 and 57.5%/54.8% ASR-LLM on Natural Questions (NQ, 361K-passage subset; not the standard 21M DPR corpus) and MS MARCO (8.8M passages), while maintaining near-benign perplexity (32.4 vs. 28.4). Cross-model evaluation across four target LLMs shows nontrivial effectiveness under a fixed CATG generator (48.6-57.5% ASR-LLM). Surrogate-transfer evaluation against unseen retrievers, including ColBERT and rebuilt indexes using commercial embedding models, yields 64.7% average HR@10 under the same injected-corpus protocol. In a sampled large-corpus evaluation built from a Wikipedia-scale 21M-passage construction, SilentRetrieval retains 74.2% HR@10 at a 0.016% poisoning ratio, characterizing large-corpus behavior under the sampled protocol. Combined retrieval-side and generation-side defenses reduce ASR-LLM to 25.6% at a 6x latency trade-off in our evaluated setting, and to 21.3% under the strongest evaluated configuration; adaptive attacks recover 6.2% HR@10 in the matched MiniLM-L6-v2 reranker setting. Human evaluation (n=600 documents, Krippendorff's α=0.74) shows substantially lower flag rates than disfluent baselines, while remaining numerically more suspicious than benign content at the current sample size (p≈0.064).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4b8e92f-58a9-4213-9c8c-a955f3ce09acBuilds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsJian Xie, Kai Zhang, Jiangjie Chen, Renze Lou et al.ICLR 2024 · 294 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- Reranker Helps, but Not Enough: Towards Strong Poisoning Attacks Against Retrieval-Augmented GenerationXiaokun Yang, Jian Liang, Yesheng Liu, Xin Xiong et al.ICML 2026
- WARP: A Word-Level Backdoor Attack Targeting RAG Systems via Retrieval Corpus PoisoningHui Liu, Yibo Zhou, Liguo Dong, Weidong Li et al.KDD 2026
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 6 citations
- Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation SystemsHaowei Wang, Rupeng Zhang, Junjie Wang, Mingyang Li et al.AAAI 2026 · 3 citations
- On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application DomainsXun Xian, Ganghua Wang, Xuan Bi, Rui Zhang et al.ICML 2025
