GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
Yueqi Xie, Minghong Fang, Renjie Pi, Neil Gong
Abstract
Large Language Models (LLMs) face threats from jailbreak prompts. Existing methods for detecting jailbreak prompts are primarily online moderation APIs or finetuned LLMs. These strategies, however, often require extensive and resource-intensive data collection and training processes. In this study, we propose GradSafe, which effectively detects jailbreak prompts by scrutinizing the gradients of safety-critical parameters in LLMs. Our method is grounded in a pivotal observation: the gradients of an LLM's loss for jailbreak prompts paired with compliance response exhibit similar patterns on certain safety-critical parameters. In contrast, safe prompts lead to different gradient patterns. Building on this observation, GradSafe analyzes the gradients from prompts (paired with compliance responses) to accurately detect jailbreak prompts. We show that GradSafe, applied to Llama-2 without further training, outperforms Llama Guard-despite its extensive finetuning with a large dataset-in detecting jailbreak prompts. This superior performance is consistent across both zero-shot and adaptation scenarios, as evidenced by our evaluations on ToxicChat and XSTest. The source code is available at https://github.com/xyq7/GradSafe .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eac9bad4-a290-4907-84ff-3222e6b768c1Cited by top-tier papers27
- Sok: Evaluating Jailbreak Guardrails for Large Language ModelsXunguang Wang, Zhenlan Ji, Wenxuan Wang, Zongjie Li et al.S&P 2026 · 27 citations
- MLLM-Protector: Ensuring MLLM's Safety without Hurting PerformanceRenjie Pi, Tianyang Han, Jianshu Zhang, Yueqi Xie et al.EMNLP 2024 · 21 citations
- JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language ModelsZifan Peng, Yule Liu, Zhen Sun, Mingchen Li et al.ICLR 2026 · 20 citations
- Monitoring Decomposition Attacks with Lightweight Sequential MonitorsYueh-Han Chen, Nitish Joshi, Yulin Chen, Maksym Andriushchenko et al.ICLR 2026 · 15 citations
- Detecting and Filtering Unsafe Training Data via Data Attribution with Denoised RepresentationYijun Pan, Taiwei Shi, Jieyu Zhao, Jiaqi MaICML 2026 · 10 citations
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami et al.NeurIPS 2020 · 1,022 citations
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
- Benchmarking and Defending against Indirect Prompt Injection Attacks on Large Language ModelsJingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman et al.KDD 2025 · 27 citations
Related papers
- GraphShield: Graph-Theoretic Modeling of Network-Level Dynamics for Robust Jailbreak DetectionSunghee Dong, Sungwon Yi, Kangmin Bae, Jaeyoon Kim et al.ICLR 2026
- Efficient Detection of Toxic Prompts in Large Language ModelsYi Liu, Junzhe Yu, Huijia Sun, Ling Shi et al.ASE 2024 · 6 citations
- Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss LandscapesXiaomeng Hu, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2024 · 97 citations
- PARDEN, Can You Repeat That? Defending against Jailbreaks via RepetitionZiyang Zhang, Qizhen Zhang, Jakob Nicolaus FoersterICML 2024 · 36 citations
- JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and ManipulationShenyi Zhang, Yuchen Zhai, Keyan Guo, Hongxin Hu et al.USENIX Security 2025
