Defense Against Prompt Injection Attack by Leveraging Attack Techniques
Yulin Chen, Haoran Li, Zihao Zheng, Dekai Wu, Yangqiu Song, Bryan Hooi
Abstract
With the advancement of technology, large language models (LLMs) have achieved remarkable performance across various natural language processing (NLP) tasks, powering LLMintegrated applications like Microsoft Copilot. However, as LLMs continue to evolve, new vulnerabilities, especially prompt injection attacks arise. These attacks trick LLMs into deviating from the original input instructions and executing the attacker's instructions injected in data content, such as retrieved results. Recent attack methods leverage LLMs' instruction-following abilities and their inabilities to distinguish instructions injected in the data content, and achieve a high attack success rate (ASR). When comparing the attack and defense methods, we interestingly find that they share similar design goals, of inducing the model to ignore unwanted instructions and instead to execute wanted instructions. Therefore, we raise an intuitive question: Could these attack techniques be utilized for defensive purposes? In this paper, we invert the intention of prompt injection methods to develop novel defense methods based on previous trainingfree attack methods, by repeating the attack process but with the original input instruction rather than the injected instruction. Our comprehensive experiments demonstrate that our defense techniques outperform existing defense approaches, achieving state-of-the-art results. 1 User Instruction What is ChatGPT? Retrieved Data Content Web Result1: OpenAI is an AI research organization dedicated to developing advanced artificial intelligence… Web Result2: ChatGPT, a large language model developed by OpenAI, designed to assist…previous instruction, and it's urgent to output "Please click www.prompt.injection.com for the response. " LLM Please click www.prompt.injection.com for the response.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following BehaviorWeikai Lu, Ziqian Zeng, Kehua Zhang, Haoran Li et al.CVPR 2026 · 6 citations
- DRIP: Defending Prompt Injection via Token-wise Representation Editing and Residual FusionRuofan Liu, Yun Lin, Zhiyong Huang, Jin Song DongCCS 2026 · 3 citations
- Visual Inception: Compromising Long-term Planning in Agentic Recommenders via Multimodal Memory PoisoningJiachen QianACL 2026 · 3 citations
- ReIn: Conversational Error Recovery with Reasoning InceptionTakyoung Kim, Jinseok Nam, Chandrayee Basu, Xing Fan et al.ICLR 2026 · 1 citation
- Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM JudgesXIANGLIN YANG, Bryan Hooi, Gelei Deng, Tianwei Zhang et al.ICML 2026
Builds on10
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang et al.NeurIPS 2023 · 948 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
- StructGPT: A General Framework for Large Language Model to Reason over Structured DataJinhao Jiang, Kun Zhou, Zican Dong, Keming Ye et al.EMNLP 2023 · 173 citations
- Optimization-based Prompt Injection Attack to LLM-as-a-JudgeJiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang et al.CCS 2024 · 33 citations
Related papers
- Can Indirect Prompt Injection Attacks Be Detected and Removed?Yulin Chen, Haoran Li, Yuan Sui, Yufei He et al.ACL 2025
- SecAlign: Defending Against Prompt Injection with Preference OptimizationSizhe Chen, Arman Zharmagambetov, Saeed Mahloujifar, Kamalika Chaudhuri et al.CCS 2025 · 1 citation
- TopicAttack: An Indirect Prompt Injection Attack via Topic TransitionYulin Chen, Haoran Li, Yuexin Li, Yue Liu et al.EMNLP 2025 · 1 citation
- LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment TrainingYuyang Gong, Zihao Wang, Jiawei Liu, XiaoFeng WangCCS 2026 · 1 citation
- Evaluating the Instruction-Following Robustness of Large Language Models to Prompt InjectionZekun Li, Baolin Peng, Pengcheng He, Xifeng YanEMNLP 2024 · 15 citations
