DRIP: Defending Prompt Injection via Token-wise Representation Editing and Residual Fusion
Ruofan Liu, Yun Lin, Zhiyong Huang, Jin Song Dong
Abstract
We anticipate that large language models (LLMs) will become deeply integrated into IT infrastructures by processing user data according to predefined instructions. However, conventional LLMs remain vulnerable to prompt injection attacks, where malicious users inject directive tokens within the data to manipulate model behavior. Leading defense strategies attempt to train LLMs to semantically distinguish between data and instruction tokens. Nevertheless, these approaches still face two key challenges: (1) maintaining a balance between utility and security, and (2) preventing the model from interpreting instruction-like semantics in the data as higherpriority directives than the intended instructions. In this work, we propose DRIP which aims to (1) precisely remove the instruction semantics from the tokens in the data section while preserving their data semantics and (2) robustly maintain the effectiveness of the intended instruction, even in the presence of strong adversarial content within the data. As for "de-instructionalize" data tokens, we propose a training paradigm across data curation, model architecture, and loss design. This paradigm introduces a lightweight representationediting module, which is trained to edit the embedding of instruction-like tokens in the data section, enhancing the model's security without compromising utility. As for the "non-overwritability" of the intended instruction, we introduce a minimal residual module in LLM to substantially reduce the ability of adversarial data content to overwrite the original instruction. We extensively compare DRIP with state-of-the-art techniques, including StruQ, SecAlign, ISE, and PFT on LLaMA-8B and Mistral-7B across three prompt injection benchmarks (SEP, AlpacaFarm, and InjecAgent). The results show that DRIP (1) improves role separation score by 12-49% and reduces attack success rate by over 66% for adaptive attacks and (2) achieves utility on par with the undefended model, indicating a new state-of-the-art against prompt injection attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang et al.NeurIPS 2023 · 948 citations
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
- Prompt Injection Attack to Tool Selection in LLM AgentsJiawen Shi, Zenghui Yuan, Guiyao Tie, Pan Zhou et al.NDSS 2026 · 181 citations
Related papers
- StruQ: Defending Against Prompt Injection with Structured QueriesSizhe Chen, Julien Piet, Chawin Sitawarin, David A. WagnerUSENIX Security 2025
- Instructional Segment Embedding: Improving LLM Safety with Instruction HierarchyTong Wu, Shujian Zhang, Kaiqiang Song, Silei Xu et al.ICLR 2025
- ASIDE: Architectural Separation of Instructions and Data in Language ModelsEgor Zverev, Evgenii Kortukov, Alexander Panfilov, Alexandra Volkova et al.ICLR 2026 · 28 citations
- LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment TrainingYuyang Gong, Zihao Wang, Jiawei Liu, XiaoFeng WangCCS 2026 · 1 citation
- Can Indirect Prompt Injection Attacks Be Detected and Removed?Yulin Chen, Haoran Li, Yuan Sui, Yufei He et al.ACL 2025
