Prompt Perturbation in Retrieval-Augmented Generation based Large Language Models
Zhibo Hu, Chen Wang, Yanfeng Shu, Hye-Young Paik, Liming Zhu
摘要
The robustness of large language models (LLMs) becomes increasingly important as their use rapidly grows in a wide range of domains. Retrieval-Augmented Generation (RAG) is considered as a means to improve the trustworthiness of text generation from LLMs. However, how the outputs from RAG-based LLMs are affected by slightly different inputs is not well studied. In this work, we find that the insertion of even a short prefix to the prompt leads to the generation of outputs far away from factually correct answers. We systematically evaluate the effect of such prefixes on RAG by introducing a novel optimization technique called Gradient Guided Prompt Perturbation (GGPP). GGPP achieves a high success rate in steering outputs of RAG-based LLMs to targeted wrong answers. It can also cope with instructions in the prompts requesting to ignore irrelevant context. We also exploit LLMs' neuron activation difference between prompts with and without GGPP perturbations to give a method that improves the robustness of RAG-based LLMs through a highly effective detector trained on neuron activation triggered by GGPP generated prompts. Our evaluation on open-sourced LLMs demonstrates the effectiveness of our methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 被引用 6 次
- Uncovering Competing Poisoning Attacks in Retrieval-Augmented GenerationLiuji Chen, Xiaofang Yang, Yuanzhuo Lu, Jinghao Zhang 等KDD 2026 · 被引用 4 次
- One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via HubnessHiroyuki Deguchi, Katsuki Chousa, Yusuke SakaiACL 2026
- AGRAG: Advanced Graph-Based Retrieval-Augmented Generation for LLMsYubo Wang, Haoyang Li, Fei Teng, Lei ChenICDE 2026
- Open Schrödinger's Closed Box: Identifying Retrieval Augmented Generation in API-Accessible Large Language Model ServicesYukun Jiang, Xinyue Shen, Michael Backes, Zheng Li 等ACL 2026
它引用的顶会 Paper16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
相关 Paper
- AIP: Subverting Retrieval-Augmented Generation via Adversarial Instructional PromptSaket S. Chaturvedi, Gaurav Bagwe, Lan Zhang, Xiaoyong YuanEMNLP 2025
- ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-SearchZeyu Shen, Basileal Imana, Tong Wu, Chong Xiang 等NeurIPS 2025 · 被引用 26 次
- Robust Fine-tuning for Retrieval Augmented Generation against Retrieval DefectsYiteng Tu, Weihang Su, Yujia Zhou, Yiqun Liu 等SIGIR 2025 · 被引用 9 次
- EmoRAG: Evaluating RAG Robustness to Symbolic PerturbationsXinyun Zhou, Xinfeng Li, Yinan Peng, Ming Xu 等KDD 2026 · 被引用 2 次
- In-depth Analysis of Graph-based RAG in a Unified FrameworkYingli Zhou, Yaodong Su, Youran Sun, Shu Wang 等VLDB 2025 · 被引用 48 次
