Defenses Against Prompt Attacks Learn Surface Heuristics
Shawn Li, Chenxiao Yu, Zhiyu Ni, Hao Li, Charith Peris, Chaowei Xiao, Yue Zhao
摘要
Large language models (LLMs) are increasingly deployed in security-sensitive applications, where they must follow system- or developer-specified instructions that define the intended task behavior, while completing benign user requests. When adversarial instructions appear in user queries or externally retrieved content, models may override intended logic. Recent defenses rely on supervised fine-tuning with benign and malicious labels. Although these methods achieve high attack rejection rates, we find that they rely on narrow correlations in defense data rather than harmful intent, leading to systematic rejection of safe inputs. We analyze three recurring shortcut behaviors induced by defense fine-tuning. Position bias arises when benign content placed later in a prompt is rejected at much higher rates; across reasoning benchmarks, suffix-task rejection rises from below 10% to as high as 90%. Token trigger bias occurs when strings common in attack data raise rejection probability even in benign contexts; inserting a single trigger token increases false refusals by up to 50%. Topic generalization bias reflects poor generalization beyond the defense data distribution, with defended models suffering test-time accuracy drops of up to 40%. These findings suggest that current prompt-injection defenses frequently respond to attack-like surface patterns rather than the underlying intent. We introduce controlled diagnostic datasets and a systematic evaluation across two base models and multiple defense pipelines, highlighting limitations of supervised fine-tuning for reliable LLM security.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Prompt Injection as Role ConfusionCharles Ye, Jasmine Cui, Dylan Hadfield-MenellICML 2026 · 被引用 6 次
- ``Someone Hid It!'': Query-Agnostic Black-Box Attacks on LLM-Based RetrievalJiate Li, Defu Cao, Li Li, Wei Yang 等ICML 2026 · 被引用 4 次
它引用的顶会 Paper6
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia 等USENIX Security 2024 · 被引用 308 次
- Benchmarking and Defending against Indirect Prompt Injection Attacks on Large Language ModelsJingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman 等KDD 2025 · 被引用 27 次
- SecAlign: Defending Against Prompt Injection with Preference OptimizationSizhe Chen, Arman Zharmagambetov, Saeed Mahloujifar, Kamalika Chaudhuri 等CCS 2025 · 被引用 1 次
- Instructional Segment Embedding: Improving LLM Safety with Instruction HierarchyTong Wu, Shujian Zhang, Kaiqiang Song, Silei Xu 等ICLR 2025
相关 Paper
- LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment TrainingYuyang Gong, Zihao Wang, Jiawei Liu, XiaoFeng WangCCS 2026 · 被引用 1 次
- The Illusion of Role Separation: Hidden Shortcuts in LLM Role Learning (and How to Fix Them)Zihao Wang, Yibo Jiang, Jiahao Yu, Heqing HuangICML 2025
- When Style Breaks Safety: Defending LLMs Against Superficial Style AlignmentYuxin Xiao, Sana Tonekaboni, Walter Gerych, Vinith Menon Suriyakumar 等ICLR 2026 · 被引用 8 次
- Fight Back Against Jailbreaking via Prompt Adversarial TuningYichuan Mo, Yuji Wang, Zeming Wei, Yisen WangNeurIPS 2024 · 被引用 90 次
- No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety MechanismsJoshua Kazdan, Abhay Puri, Rylan Schaeffer, Lisa Yu 等ICLR 2026
