Lune

CCS2026顶会

DRIP: Defending Prompt Injection via Token-wise Representation Editing and Residual Fusion

Ruofan Liu, Yun Lin, Zhiyong Huang, Jin Song Dong

2026年份
3被引次数

摘要

We anticipate that large language models (LLMs) will become deeply integrated into IT infrastructures by processing user data according to predefined instructions. However, conventional LLMs remain vulnerable to prompt injection attacks, where malicious users inject directive tokens within the data to manipulate model behavior. Leading defense strategies attempt to train LLMs to semantically distinguish between data and instruction tokens. Nevertheless, these approaches still face two key challenges: (1) maintaining a balance between utility and security, and (2) preventing the model from interpreting instruction-like semantics in the data as higherpriority directives than the intended instructions. In this work, we propose DRIP which aims to (1) precisely remove the instruction semantics from the tokens in the data section while preserving their data semantics and (2) robustly maintain the effectiveness of the intended instruction, even in the presence of strong adversarial content within the data. As for "de-instructionalize" data tokens, we propose a training paradigm across data curation, model architecture, and loss design. This paradigm introduces a lightweight representationediting module, which is trained to edit the embedding of instruction-like tokens in the data section, enhancing the model's security without compromising utility. As for the "non-overwritability" of the intended instruction, we introduce a minimal residual module in LLM to substantially reduce the ability of adversarial data content to overwrite the original instruction. We extensively compare DRIP with state-of-the-art techniques, including StruQ, SecAlign, ISE, and PFT on LLaMA-8B and Mistral-7B across three prompt injection benchmarks (SEP, AlpacaFarm, and InjecAgent). The results show that DRIP (1) improves role separation score by 12-49% and reduces attack success rate by over 66% for adaptive attacks and (2) achieves utility on par with the undefended model, indicating a new state-of-the-art against prompt injection attacks.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper19

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖