SoK: Robustness in Large Language Models against Jailbreak Attacks
Feiyue Xu, Hongsheng Hu, Chaoxiang He, Sheng Hang, Hanqing Hu, Xiuming Liu, Yubo Zhao, Zhengyan Zhou, Bin Benjamin Zhu, Shifeng Sun, Dawu Gu, Shuo Wang
Abstract
Large Language Models (LLMs) have achieved remarkable success but remain highly susceptible to jailbreak attacks, where adversarial prompts that coerce models into generating harmful, unethical, or policy-violating outputs. Such attacks pose real-world risks, eroding safety, trust, and regulatory compliance in high-stakes applications. Although a variety of attack and defense methods have been proposed, existing evaluation practices are inadequate, often relying on narrow metrics like attack success rate that fail to capture the multidimensional nature of LLM security. In this paper, we present a systematic taxonomy of jailbreak attacks and defenses and introduce Security Cube, a unified, multidimensional framework for comprehensive evaluation of these techniques. We provide detailed comparison tables of existing attacks and defenses, highlighting key insights and open challenges across the literature. Leveraging Security Cube, we conduct benchmark studies on 13 representative attacks and 5 defenses, establishing a clear view of the current landscape encompassing jailbreak attacks, defenses, automated judges, and LLM vulnerabilities. Based on these evaluations, we distill critical findings, identify unresolved problems, and outline promising research directions for enhancing LLM robustness against jailbreak attacks. Our analysis aims to pave the way towards more robust, interpretable, and trustworthy LLM systems. Our code is available at Code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba804b81-9558-4136-bc21-0992eb6f6a65Builds on34
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
Related papers
- From LLMs to MLLMs: Exploring the Landscape of Multimodal JailbreakingSiyuan Wang, Zhuohan Long, Zhihao Fan, Zhongyu WeiEMNLP 2024 · 5 citations
- Sok: Evaluating Jailbreak Guardrails for Large Language ModelsXunguang Wang, Zhenlan Ji, Wenxuan Wang, Zongjie Li et al.S&P 2026 · 27 citations
- The Jailbreak Tax: How Useful are Your Jailbreak Outputs?Kristina Nikolic, Luze Sun, Jie Zhang, Florian TramèrICML 2025
- JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMsJunjie Chu, Yugeng Liu, Ziqing Yang, Xinyue Shen et al.ACL 2025
- MASTERKEY: Automated Jailbreaking of Large Language Model ChatbotsGelei Deng, Yi Liu, Yuekang Li, Kailong Wang et al.NDSS 2024
