PIGuard: Prompt Injection Guardrail via Mitigating Overdefense for Free
Hao Li, Xiaogeng Liu, Ning Zhang, Chaowei Xiao
摘要
Prompt injection attacks pose a critical threat to large language models (LLMs), enabling goal hijacking and data leakage. Prompt guard models, though effective in defense, suffer from over-defense—falsely flagging benign inputs as malicious due to trigger word bias. To address this issue, we introduce NotInject, an evaluation dataset that systematically measures over-defense across various prompt guard models. NotInject contains 339 benign samples enriched with trigger words common in prompt injection attacks, enabling fine-grained evaluation. Our results show that state-of-the-art models suffer from over-defense issues, with accuracy dropping close to random guessing levels (60%). To mitigate this, we propose PIGuard, a novel prompt guard model that incorporates a new training strategy, Mitigating Over-defense for Free (MOF), which significantly reduces the bias on trigger words. PIGuard demonstrates state-of-the-art performance on diverse benchmarks including Not-Inject, surpassing the existing best model by 30.4%, offering a robust and open-source so-lution for detecting prompt injection attacks. The code and datasets are released at https: //github.com/leolee99/PIGuard .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt InjectionsMilad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff 等USENIX Security 2026 · 被引用 134 次
- DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM AgentsHao Li, Xiaogeng Liu, Hung-Chun Chiu, Dianqi Li 等NeurIPS 2025 · 被引用 76 次
- AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool InvocationsYu He, Haozhe Zhu, Yiming Li, Shuo Shao 等USENIX Security 2026 · 被引用 45 次
- Bypassing Prompt Guards in Production with Controlled-Release PromptingJaiden Fairoze, Sanjam Garg, Keewoo Lee, Mingyuan WangUSENIX Security 2026 · 被引用 8 次
- PIArena: A Platform for Prompt Injection EvaluationRunpeng Geng, Chenlong Yin, Yanting Wang, Ying Chen 等ACL 2026 · 被引用 6 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia 等USENIX Security 2024 · 被引用 308 次
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen 等CCS 2024 · 被引用 132 次
- Tensor Trust: Interpretable Prompt Injection Attacks from an Online GameSam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato 等ICLR 2024 · 被引用 123 次
相关 Paper
- Can Indirect Prompt Injection Attacks Be Detected and Removed?Yulin Chen, Haoran Li, Yuan Sui, Yufei He 等ACL 2025
- Defense Against Prompt Injection Attack by Leveraging Attack TechniquesYulin Chen, Haoran Li, Zihao Zheng, Dekai Wu 等ACL 2025
- LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment TrainingYuyang Gong, Zihao Wang, Jiawei Liu, XiaoFeng WangCCS 2026 · 被引用 1 次
- SecAlign: Defending Against Prompt Injection with Preference OptimizationSizhe Chen, Arman Zharmagambetov, Saeed Mahloujifar, Kamalika Chaudhuri 等CCS 2025 · 被引用 1 次
- IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM AgentsHengyu An, Jinghuai Zhang, Tianyu Du, Chunyi Zhou 等EMNLP 2025
