Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
Kai Hu, Abhinav Aggarwal, Mehran Khodabandeh, David Zhang, Eric Hsin, Li Chen, Ankit Jain, Matt Fredrikson, Akash Bharadwaj
Abstract
This paper advances Automated Red Teaming (ART) for evaluating Large Language Model (LLM) safety through both methodological and evaluation contributions. We first analyze existing example-based red teaming approaches and identify critical limitations in scalability and validity, and propose a policy-based evaluation framework that defines harmful content through safety policies rather than examples. This framework incorporates multiple objectives beyond attack success rate (ASR), including risk coverage, semantic diversity, and fidelity to desired data distributions. We then analyze the Pareto trade-offs between these objectives. Our second contribution, Jailbreak-Zero, is a novel ART method that adapts to this evaluation framework. Jailbreak-Zero can be a zero-shot method that generates successful jailbreak prompts with minimal human input, or a finetuned method where the attack LLM explores and exploits the vulnerabilities of a particular victim to achieve Pareto-optimality. Moreover, it exposes controls to navigate Pareto trade-offs as required by a use case without re-training. Jailbreak-Zero achieves superior attack success rates with human-readable attacks compared to prior methods while maximizing semantic diversity and fidelity to any desired data distribution. Our results generalize across both open-source (Llama, Qwen, Mistral) and proprietary models (GPT-4o and Claude 3.5). Lastly, our method retains efficacy even after the LLM that we are red-teaming undergoes safety alignment to mitigate the risks exposed by a previous round of red teaming.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 630aa9fa-68b0-4ff9-a382-0e2d141209baBuilds on14
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang et al.ICLR 2024 · 441 citations
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
- Rainbow Teaming: Open-Ended Generation of Diverse Adversarial PromptsMikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro et al.NeurIPS 2024 · 231 citations
Related papers
- h4rm3l: A Language for Composable Jailbreak Attack SynthesisMoussa Koulako Bala Doumbouya, Ananjan Nandi, Gabriel Poesia, Davide Ghilardi et al.ICLR 2025
- SHARP: Self-adaptive Harmful Category-aware Prompt Generation for Black-box JailbreakingYingjie Xue, Xingyou Xia, Jun Zhang, Yunbo Cao et al.ACL 2026
- Red Teaming LLMs via Linguistic-Aware FuzzingShuai Yuan, Nian Luo, Jingling Sun, Yihao Huang et al.FSE 2026
- CoP: Agentic Red-teaming for Large Language Models using Composition of PrinciplesChen Xiong, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2025 · 13 citations
- ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign UsersGuanlin Li, Kangjie Chen, Shudong Zhang, Jie Zhang et al.NeurIPS 2024 · 39 citations
