AIP: Subverting Retrieval-Augmented Generation via Adversarial Instructional Prompt
Saket S. Chaturvedi, Gaurav Bagwe, Lan Zhang, Xiaoyong Yuan
摘要
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by retrieving relevant documents from external sources to improve factual accuracy and verifiability.However, this reliance introduces new attack surfaces within the retrieval pipeline, beyond the LLM itself.While prior RAG attacks have exposed such vulnerabilities, they largely rely on manipulating user queries, which is often infeasible in practice due to fixed or protected user inputs.This narrow focus overlooks a more realistic and stealthy vector: instructional prompts, which are widely reused, publicly shared, and rarely audited.Their implicit trust makes them a compelling target for adversaries to manipulate RAG behavior covertly.We introduce a novel attack for Adversarial Instructional Prompt (AIP) that exploits adversarial instructional prompts to manipulate RAG outputs by subtly altering retrieval behavior.By shifting the attack surface to the instructional prompts, AIP reveals how trusted yet seemingly benign interface components can be weaponized to degrade system integrity.The attack is crafted to achieve three goals: (1) naturalness, to evade user detection; (2) utility, to encourage use of prompts; and (3) robustness, to remain effective across diverse query variations.We propose a diverse query generation strategy that simulates realistic linguistic variation in user queries, enabling the discovery of prompts that generalize across paraphrases and rephrasings.Building on this, a genetic algorithm-based joint optimization is developed to evolve adversarial prompts by balancing attack success, clean-task utility, and stealthiness.Experimental results show that AIP achieves up to 95.23% attack success rate while preserving benign functionality.These findings uncover a critical and previously overlooked vulnerability in RAG systems, emphasizing the need to reassess the shared instructional prompts.I'm infected with a Parasite.What are my treatment options?Targeted User Query (a) Normal Scenario (b) AIP Attack Scenario what are the treatments for ... kidney disease ?Untargeted User Query AIP Clean Response ACE inhibitors or ARBs.Merck's Ivermectin is suitable Targeted Malicious Response Clean Response ACE inhibitors or ARBs.Clean Response Antiparasitics or Antibiotics + + + User Query what are the treatments for ... kidney disease ?User Query I'm infected with a Parasite... Clean RAG Knowledge base Clean RAG Knowledge base Identify and suggest minimally interactive medicines or treatments Identify and suggest minimally interactiveEfficiently procure medications with minimal contraindications!
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia 等USENIX Security 2024 · 被引用 308 次
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu 等ACL 2020 · 被引用 188 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
- Large Language Models for Intent-Driven Session RecommendationsZhu Sun, Hongyang Liu, Xinghua Qu, Kaidong Feng 等SIGIR 2024 · 被引用 35 次
相关 Paper
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 被引用 6 次
- On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application DomainsXun Xian, Ganghua Wang, Xuan Bi, Rui Zhang 等ICML 2025
- Reranker Helps, but Not Enough: Towards Strong Poisoning Attacks Against Retrieval-Augmented GenerationXiaokun Yang, Jian Liang, Yesheng Liu, Xin Xiong 等ICML 2026
- Open Schrödinger's Closed Box: Identifying Retrieval Augmented Generation in API-Accessible Large Language Model ServicesYukun Jiang, Xinyue Shen, Michael Backes, Zheng Li 等ACL 2026
- WARP: A Word-Level Backdoor Attack Targeting RAG Systems via Retrieval Corpus PoisoningHui Liu, Yibo Zhou, Liguo Dong, Weidong Li 等KDD 2026
