DRIP: Defending Prompt Injection via Token-wise Representation Editing and Residual Fusion
Ruofan Liu, Yun Lin, Zhiyong Huang, Jin Song Dong
摘要
We anticipate that large language models (LLMs) will become deeply integrated into IT infrastructures by processing user data according to predefined instructions. However, conventional LLMs remain vulnerable to prompt injection attacks, where malicious users inject directive tokens within the data to manipulate model behavior. Leading defense strategies attempt to train LLMs to semantically distinguish between data and instruction tokens. Nevertheless, these approaches still face two key challenges: (1) maintaining a balance between utility and security, and (2) preventing the model from interpreting instruction-like semantics in the data as higherpriority directives than the intended instructions. In this work, we propose DRIP which aims to (1) precisely remove the instruction semantics from the tokens in the data section while preserving their data semantics and (2) robustly maintain the effectiveness of the intended instruction, even in the presence of strong adversarial content within the data. As for "de-instructionalize" data tokens, we propose a training paradigm across data curation, model architecture, and loss design. This paradigm introduces a lightweight representationediting module, which is trained to edit the embedding of instruction-like tokens in the data section, enhancing the model's security without compromising utility. As for the "non-overwritability" of the intended instruction, we introduce a minimal residual module in LLM to substantially reduce the ability of adversarial data content to overwrite the original instruction. We extensively compare DRIP with state-of-the-art techniques, including StruQ, SecAlign, ISE, and PFT on LLaMA-8B and Mistral-7B across three prompt injection benchmarks (SEP, AlpacaFarm, and InjecAgent). The results show that DRIP (1) improves role separation score by 12-49% and reduces attack success rate by over 66% for adaptive attacks and (2) achieves utility on par with the undefended model, indicating a new state-of-the-art against prompt injection attacks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia 等USENIX Security 2024 · 被引用 308 次
- Prompt Injection Attack to Tool Selection in LLM AgentsJiawen Shi, Zenghui Yuan, Guiyao Tie, Pan Zhou 等NDSS 2026 · 被引用 181 次
相关 Paper
- StruQ: Defending Against Prompt Injection with Structured QueriesSizhe Chen, Julien Piet, Chawin Sitawarin, David A. WagnerUSENIX Security 2025
- Instructional Segment Embedding: Improving LLM Safety with Instruction HierarchyTong Wu, Shujian Zhang, Kaiqiang Song, Silei Xu 等ICLR 2025
- ASIDE: Architectural Separation of Instructions and Data in Language ModelsEgor Zverev, Evgenii Kortukov, Alexander Panfilov, Alexandra Volkova 等ICLR 2026 · 被引用 28 次
- LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment TrainingYuyang Gong, Zihao Wang, Jiawei Liu, XiaoFeng WangCCS 2026 · 被引用 1 次
- Can Indirect Prompt Injection Attacks Be Detected and Removed?Yulin Chen, Haoran Li, Yuan Sui, Yufei He 等ACL 2025
