USENIX Security2026Top-tier venue
Confundo: Learning to Generate Robust Poison for Practical RAG Systems
Haoyang Hu, Zhejun Jiang, Yueming Lyu, Junyuan Zhang, Yi Liu, Ka-Ho Chow
Abstract
Retrieval-augmented generation (RAG) is increasingly deployed in real-world applications, where its referencegrounded design makes outputs appear trustworthy. This trust has spurred research on poisoning attacks that craft malicious content, inject it into knowledge sources, and manipulate RAG responses. However, when evaluated in practical RAG systems, existing attacks suffer from severely degraded effectiveness. This gap stems from two overlooked realities: (i) content is often processed before use, which can fragment the poison and weaken its effect, and (ii) users often do not issue the exact queries anticipated during attack design. These factors can lead practitioners to underestimate risks and develop a false sense of security. To better characterize the threat to practical systems, we present Confundo 1 , a learning-to-poison framework that fine-tunes a large language model as a poison generator to achieve high effectiveness, robustness, and stealthiness. Confundo provides a unified framework supporting multiple attack objectives, demonstrated by manipulating factual correctness, inducing biased opinions, and triggering hallucinations. By addressing these overlooked challenges, Confundo consistently outperforms a wide range of purposebuilt attacks across datasets and RAG configurations by large margins, even in the presence of defenses. Beyond exposing vulnerabilities, we also present a defensive use case that protects web content from unauthorized incorporation into RAG systems via scraping, with no impact on user experience.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d170c56a-c47e-4f98-a275-c73b2aab32e8Builds on27
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- INSIDE: LLMs' Internal States Retain the Power of Hallucination DetectionChao Chen, Kai Liu, Ze Chen, Yi Gu et al.ICLR 2024 · 281 citations
- MedRAG: Enhancing Retrieval-augmented Generation with Knowledge Graph-Elicited Reasoning for Healthcare CopilotXuejiao Zhao, Siyan Liu, Su-Yin Yang, Chunyan MiaoWWW 2025 · 134 citations
- HaloScope: Harnessing Unlabeled LLM Generations for Hallucination DetectionXuefeng Du, Chaowei Xiao, Sharon LiNeurIPS 2024 · 131 citations
- Fast Adversarial Attacks on Language Models In One GPU MinuteVinu Sankar Sadasivan, Shoumik Saha, Gaurang Sriramanan, Priyatham Kattakinda et al.ICML 2024 · 85 citations
Related papers
- MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Knowledge Poisoning AttacksHyeonjeong Ha, Qiusi Zhan, Jeonghwan Kim, Dimitrios Bralios et al.ACL 2026 · 1 citation
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 6 citations
- Reranker Helps, but Not Enough: Towards Strong Poisoning Attacks Against Retrieval-Augmented GenerationXiaokun Yang, Jian Liang, Yesheng Liu, Xin Xiong et al.ICML 2026
- DisarmRAG: Stealthy Retriever Poisoning to Disable Self-Correction in Retrieval-Augmented GenerationYanbo Dai, Zhenlan Ji), Zongjie Li, Kuan Li et al.CCS 2026 · 2 citations
- Uncovering Competing Poisoning Attacks in Retrieval-Augmented GenerationLiuji Chen, Xiaofang Yang, Yuanzhuo Lu, Jinghao Zhang et al.KDD 2026 · 4 citations
