CachePrune: Teaching LLMs What Not to Follow via KV-Cache Editing
Rui Wang, Junda Wu, Yu Xia, Tong Yu, Ruiyi Zhang, Ryan A. Rossi, Subrata Mitra, Lina Yao, Julian J. McAuley
摘要
Large Language Models (LLMs) are susceptible to indirect prompt injection attacks, where the model inadvertently responds to instructions injected into the prompt context. This vulnerability stems from LLMs'inability to distinguish between data and instructions within a prompt. We propose CachePrune, which defends against this attack by identifying and pruning neurons associated with instruction-following during KV cache encoding of the prompt context. The pruning steers the LLM toward interpreting the context purely as data rather than as instructions to follow. To identify these neurons, we introduce a neural attribution mechanism guided by a preferential attribution loss, and theoretically connect this loss to an upper bound of the Direct Preference Optimization (DPO) objective. Further, we improve the fidelity of neural attribution by leveraging an observed triggering effect in instruction-following. Our approach does not interfere with prompt formatting or incur test-time overhead during response generation. Experiments show that CachePrune significantly reduces the attack success rate while preserving the LLM's ability to follow user instructions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- Transformer Feed-Forward Layers Are Key-Value MemoriesMor Geva, Roei Schuster, Jonathan Berant, Omer LevyEMNLP 2021 · 被引用 33 次
- PIGuard: Prompt Injection Guardrail via Mitigating Overdefense for FreeHao Li, Xiaogeng Liu, Ning Zhang, Chaowei XiaoACL 2025 · 被引用 33 次
- Benchmarking and Defending against Indirect Prompt Injection Attacks on Large Language ModelsJingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman 等KDD 2025 · 被引用 27 次
相关 Paper
- TopicAttack: An Indirect Prompt Injection Attack via Topic TransitionYulin Chen, Haoran Li, Yuexin Li, Yue Liu 等EMNLP 2025 · 被引用 1 次
- Evaluating the Instruction-Following Robustness of Large Language Models to Prompt InjectionZekun Li, Baolin Peng, Pengcheng He, Xifeng YanEMNLP 2024 · 被引用 15 次
- RedVisor: Reasoning-Aware Prompt Injection Defense via Zero-Copy KV Cache ReuseMingrui Liu, Sixiao Zhang, Cheng Long, Kwok Yan LamICML 2026 · 被引用 2 次
- Can Indirect Prompt Injection Attacks Be Detected and Removed?Yulin Chen, Haoran Li, Yuan Sui, Yufei He 等ACL 2025
- Localize and Neutralize: Gradient-Guided Token Suppression Against Visual Prompt Injection AttackDongpeng Zhang, Ke Ma, Yangbangyan Jiang, Gaozheng Pei 等ICML 2026
