AIP: Subverting Retrieval-Augmented Generation via Adversarial Instructional Prompt
Saket S. Chaturvedi, Gaurav Bagwe, Lan Zhang, Xiaoyong Yuan
Abstract
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by retrieving relevant documents from external sources to improve factual accuracy and verifiability.However, this reliance introduces new attack surfaces within the retrieval pipeline, beyond the LLM itself.While prior RAG attacks have exposed such vulnerabilities, they largely rely on manipulating user queries, which is often infeasible in practice due to fixed or protected user inputs.This narrow focus overlooks a more realistic and stealthy vector: instructional prompts, which are widely reused, publicly shared, and rarely audited.Their implicit trust makes them a compelling target for adversaries to manipulate RAG behavior covertly.We introduce a novel attack for Adversarial Instructional Prompt (AIP) that exploits adversarial instructional prompts to manipulate RAG outputs by subtly altering retrieval behavior.By shifting the attack surface to the instructional prompts, AIP reveals how trusted yet seemingly benign interface components can be weaponized to degrade system integrity.The attack is crafted to achieve three goals: (1) naturalness, to evade user detection; (2) utility, to encourage use of prompts; and (3) robustness, to remain effective across diverse query variations.We propose a diverse query generation strategy that simulates realistic linguistic variation in user queries, enabling the discovery of prompts that generalize across paraphrases and rephrasings.Building on this, a genetic algorithm-based joint optimization is developed to evolve adversarial prompts by balancing attack success, clean-task utility, and stealthiness.Experimental results show that AIP achieves up to 95.23% attack success rate while preserving benign functionality.These findings uncover a critical and previously overlooked vulnerability in RAG systems, emphasizing the need to reassess the shared instructional prompts.I'm infected with a Parasite.What are my treatment options?Targeted User Query (a) Normal Scenario (b) AIP Attack Scenario what are the treatments for ... kidney disease ?Untargeted User Query AIP Clean Response ACE inhibitors or ARBs.Merck's Ivermectin is suitable Targeted Malicious Response Clean Response ACE inhibitors or ARBs.Clean Response Antiparasitics or Antibiotics + + + User Query what are the treatments for ... kidney disease ?User Query I'm infected with a Parasite... Clean RAG Knowledge base Clean RAG Knowledge base Identify and suggest minimally interactive medicines or treatments Identify and suggest minimally interactiveEfficiently procure medications with minimal contraindications!
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8707a58b-ddc0-4069-a071-c768be8b18d5Builds on5
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu et al.ACL 2020 · 188 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Large Language Models for Intent-Driven Session RecommendationsZhu Sun, Hongyang Liu, Xinghua Qu, Kaidong Feng et al.SIGIR 2024 · 35 citations
Related papers
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 6 citations
- On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application DomainsXun Xian, Ganghua Wang, Xuan Bi, Rui Zhang et al.ICML 2025
- Reranker Helps, but Not Enough: Towards Strong Poisoning Attacks Against Retrieval-Augmented GenerationXiaokun Yang, Jian Liang, Yesheng Liu, Xin Xiong et al.ICML 2026
- Open Schrödinger's Closed Box: Identifying Retrieval Augmented Generation in API-Accessible Large Language Model ServicesYukun Jiang, Xinyue Shen, Michael Backes, Zheng Li et al.ACL 2026
- WARP: A Word-Level Backdoor Attack Targeting RAG Systems via Retrieval Corpus PoisoningHui Liu, Yibo Zhou, Liguo Dong, Weidong Li et al.KDD 2026
