FlippedRAG: Black-Box Opinion Manipulation Adversarial Attacks to Retrieval-Augmented Generation Models
Zhuo Chen, Yuyang Gong, Jiawei Liu, Miaokun Chen, Haotan Liu, Qikai Cheng, Fan Zhang, Wei Lu, Xiaozhong Liu
Abstract
Retrieval-Augmented Generation (RAG) enriches LLMs by dynamically retrieving external knowledge, reducing hallucinations and satisfying real-time information needs. While existing research mainly targets RAG's performance and efficiency, emerging studies highlight critical security concerns. Yet, current adversarial approaches remain limited, mostly addressing white-box scenarios or heuristic black-box attacks without fully investigating vulnerabilities in the retrieval phase. Additionally, prior works mainly focus on factoid Q&A tasks, their attacks lack complexity and can be easily corrected by advanced LLMs. In this paper, we investigate a more realistic and critical threat scenario: adversarial attacks intended for opinion manipulation against black-box RAG models, particularly on controversial topics. Specifically, we propose FlippedRAG, a transfer-based adversarial attack against black-box RAG-like systems. We first demonstrate that the underlying retriever of a black-box RAG can be reverse-engineered and approximated by enumerating critical queries, candidates, and answers, enabling us to train a surrogate retriever. Leveraging the surrogate retriever, we further craft target poisoning triggers, altering vary few documents to effectively manipulate both retrieval and subsequent generation, transferring the attack to the original black-box RAG model. Extensive empirical results show that FlippedRAG substantially outperforms baseline methods, improving the average attack success rate by 16.7%. Across four diverse domains, FlippedRAG achieves on average a 50% directional shift in the opinion polarity of RAG-generated responses, ultimately causing a notable 20% shift in user cognition. Furthermore, we actively evaluate the performance of several potential defensive measures, concluding that existing mitigation strategies remain insufficient against such sophisticated manipulation attacks. These results highlight an urgent need for developing innovative defensive solutions to ensure the security and trustworthiness of RAG systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2839ac55-6f07-4ba8-b724-a386aef74b51Cited by top-tier papers4
- MemPoison: Bypassing Selective Memory Mechanisms to Plant Backdoors in LLM AgentsHongtao Wang, Se Yang, Yu Chen, Puzhuo LiuCCS 2026 · 5 citations
- Confundo: Learning to Generate Robust Poison for Practical RAG SystemsHaoyang Hu, Zhejun Jiang, Yueming Lyu, Junyuan Zhang et al.USENIX Security 2026 · 5 citations
- Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented GenerationWentao Zhang, Yan Zhuang, ZhuHang Zheng, Mingfei Zhang et al.ACL 2026
- Topic-FlipRAG: Topic-Orientated Adversarial Opinion Manipulation Attacks to Retrieval-Augmented Generation ModelsYuyang Gong, Zhuo Chen, Jiawei Liu, Miaokun Chen et al.USENIX Security 2025
Builds on15
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Poisoning Web-Scale Training Datasets is PracticalNicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka et al.S&P 2024 · 309 citations
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
- (De)Randomized Smoothing for Certifiable Defense against Patch AttacksAlexander Levine, Soheil FeiziNeurIPS 2020 · 188 citations
- PatchGuard: A Provably Robust Defense against Adversarial Patches via Small Receptive Fields and MaskingChong Xiang, Arjun Nitin Bhagoji, Vikash Sehwag, Prateek MittalUSENIX Security 2021 · 172 citations
Related papers
- ``Someone Hid It!'': Query-Agnostic Black-Box Attacks on LLM-Based RetrievalJiate Li, Defu Cao, Li Li, Wei Yang et al.ICML 2026 · 4 citations
- Open Schrödinger's Closed Box: Identifying Retrieval Augmented Generation in API-Accessible Large Language Model ServicesYukun Jiang, Xinyue Shen, Michael Backes, Zheng Li et al.ACL 2026
- MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Knowledge Poisoning AttacksHyeonjeong Ha, Qiusi Zhan, Jeonghwan Kim, Dimitrios Bralios et al.ACL 2026 · 1 citation
- Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker DocumentsAvital Shafran, Roei Schuster, Vitaly ShmatikovUSENIX Security 2025
- Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation SystemsHaowei Wang, Rupeng Zhang, Junjie Wang, Mingyang Li et al.AAAI 2026 · 3 citations
