PIGuard: Prompt Injection Guardrail via Mitigating Overdefense for Free
Hao Li, Xiaogeng Liu, Ning Zhang, Chaowei Xiao
Abstract
Prompt injection attacks pose a critical threat to large language models (LLMs), enabling goal hijacking and data leakage. Prompt guard models, though effective in defense, suffer from over-defense—falsely flagging benign inputs as malicious due to trigger word bias. To address this issue, we introduce NotInject, an evaluation dataset that systematically measures over-defense across various prompt guard models. NotInject contains 339 benign samples enriched with trigger words common in prompt injection attacks, enabling fine-grained evaluation. Our results show that state-of-the-art models suffer from over-defense issues, with accuracy dropping close to random guessing levels (60%). To mitigate this, we propose PIGuard, a novel prompt guard model that incorporates a new training strategy, Mitigating Over-defense for Free (MOF), which significantly reduces the bias on trigger words. PIGuard demonstrates state-of-the-art performance on diverse benchmarks including Not-Inject, surpassing the existing best model by 30.4%, offering a robust and open-source so-lution for detecting prompt injection attacks. The code and datasets are released at https: //github.com/leolee99/PIGuard .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b0133e5-700b-489d-a2d2-2a565d5e9d4eCited by top-tier papers10
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt InjectionsMilad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff et al.USENIX Security 2026 · 134 citations
- DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM AgentsHao Li, Xiaogeng Liu, Hung-Chun Chiu, Dianqi Li et al.NeurIPS 2025 · 76 citations
- AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool InvocationsYu He, Haozhe Zhu, Yiming Li, Shuo Shao et al.USENIX Security 2026 · 45 citations
- Bypassing Prompt Guards in Production with Controlled-Release PromptingJaiden Fairoze, Sanjam Garg, Keewoo Lee, Mingyuan WangUSENIX Security 2026 · 8 citations
- PIArena: A Platform for Prompt Injection EvaluationRunpeng Geng, Chenlong Yin, Yanting Wang, Ying Chen et al.ACL 2026 · 6 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen et al.CCS 2024 · 132 citations
- Tensor Trust: Interpretable Prompt Injection Attacks from an Online GameSam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato et al.ICLR 2024 · 123 citations
Related papers
- Can Indirect Prompt Injection Attacks Be Detected and Removed?Yulin Chen, Haoran Li, Yuan Sui, Yufei He et al.ACL 2025
- Defense Against Prompt Injection Attack by Leveraging Attack TechniquesYulin Chen, Haoran Li, Zihao Zheng, Dekai Wu et al.ACL 2025
- LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment TrainingYuyang Gong, Zihao Wang, Jiawei Liu, XiaoFeng WangCCS 2026 · 1 citation
- SecAlign: Defending Against Prompt Injection with Preference OptimizationSizhe Chen, Arman Zharmagambetov, Saeed Mahloujifar, Kamalika Chaudhuri et al.CCS 2025 · 1 citation
- IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM AgentsHengyu An, Jinghuai Zhang, Tianyu Du, Chunyi Zhou et al.EMNLP 2025
