Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
Lang Gao, Jiahui Geng, Xiangliang Zhang, Preslav Nakov, Xiuying Chen
摘要
Jailbreaking in large language models (LLMs) poses major security risks by tricking models into generating harmful text. However, there is limited understanding of how jailbreaking operates, which makes it difficult to develop effective defenses. In this work, we conduct a large-scale analysis of seven jailbreak methods and uncover that inconsistencies in previous studies arise from insufficient observation samples. Our analysis reveals that jailbreaks shift harmful activations beyond a defined safety boundary, where LLMs become less sensitive to harmful information. We also find that the low and the middle layers are critical in driving these shifts, while deeper layers play a lesser role. Leveraging these insights, we propose a novel defense mechanism called Activation Boundary Defense (ABD), which adaptively constrains activations within the safety boundary. To further optimize performance, we use Bayesian optimization to select the most effective layers for ABD application and confirm that the low and middle layers have the greatest impact, consistent with our earlier observations. Experiments across multiple benchmarks demonstrate that ABD achieves an average defense success rate (DSR) of over 98% against various jailbreak attacks, with less than 2% impact on the model's overall capabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Reasoning as an Adaptive Defense for SafetyTaeyoun Kim, Fahim Tajwar, Aditi Raghunathan, Aviral KumarNeurIPS 2025 · 被引用 24 次
- Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language ModelsZirui Song, Qian Jiang, Mingxuan Cui, Mingzhe Li 等ACL 2026 · 被引用 20 次
- Mitigating the Safety Alignment Tax with Null-Space Constrained Policy OptimizationYifan Niu, Han Xiao, Dongyi Liu, Nuo Chen 等ICLR 2026 · 被引用 13 次
- Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language ModelsZixiang Xu, Yanbo Wang, Yue Huang, Xiuying Chen 等ACL 2025 · 被引用 5 次
- SoK: Robustness in Large Language Models against Jailbreak AttacksFeiyue Xu, Hongsheng Hu, Chaoxiang He, Sheng Hang 等S&P 2026 · 被引用 4 次
它引用的顶会 Paper14
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context LearningBill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri 等ICLR 2024 · 被引用 299 次
- COLD-Attack: Jailbreaking LLMs with Stealthiness and ControllabilityXingang Guo, Fangxu Yu, Huan Zhang, Lianhui Qin 等ICML 2024 · 被引用 173 次
- Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language ModelsZhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron 等USENIX Security 2024 · 被引用 103 次
- ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMsFengqing Jiang, Zhangchen Xu, Luyao Niu, Zhen Xiang 等ACL 2024 · 被引用 36 次
相关 Paper
- MirrorShield: Towards Dynamic Adaptive Defense Against Jailbreaks via Entropy-Guided Mirror CraftingRui Pu, Chaozhuo Li, Rui Ha, Litian Zhang 等AAAI 2026
- JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and ManipulationShenyi Zhang, Yuchen Zhai, Keyan Guo, Hongxin Hu 等USENIX Security 2025
- Alignment-Enhanced Decoding: Defending Jailbreaks via Token-Level Adaptive Refining of Probability DistributionsQuan Liu, Zhenhong Zhou, Longzhu He, Yi Liu 等EMNLP 2024 · 被引用 1 次
- MASTERKEY: Automated Jailbreaking of Large Language Model ChatbotsGelei Deng, Yi Liu, Yuekang Li, Kailong Wang 等NDSS 2024
- Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained OptimizationKai Hu, Weichen Yu, Yining Li, Tianjun Yao 等NeurIPS 2024 · 被引用 31 次
