Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
Kai Hu, Abhinav Aggarwal, Mehran Khodabandeh, David Zhang, Eric Hsin, Li Chen, Ankit Jain, Matt Fredrikson, Akash Bharadwaj
摘要
This paper advances Automated Red Teaming (ART) for evaluating Large Language Model (LLM) safety through both methodological and evaluation contributions. We first analyze existing example-based red teaming approaches and identify critical limitations in scalability and validity, and propose a policy-based evaluation framework that defines harmful content through safety policies rather than examples. This framework incorporates multiple objectives beyond attack success rate (ASR), including risk coverage, semantic diversity, and fidelity to desired data distributions. We then analyze the Pareto trade-offs between these objectives. Our second contribution, Jailbreak-Zero, is a novel ART method that adapts to this evaluation framework. Jailbreak-Zero can be a zero-shot method that generates successful jailbreak prompts with minimal human input, or a finetuned method where the attack LLM explores and exploits the vulnerabilities of a particular victim to achieve Pareto-optimality. Moreover, it exposes controls to navigate Pareto trade-offs as required by a use case without re-training. Jailbreak-Zero achieves superior attack success rates with human-readable attacks compared to prior methods while maximizing semantic diversity and fidelity to any desired data distribution. Our results generalize across both open-source (Llama, Qwen, Mistral) and proprietary models (GPT-4o and Claude 3.5). Lastly, our method retains efficacy even after the LLM that we are red-teaming undergoes safety alignment to mitigate the risks exposed by a previous round of red teaming.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang 等ICLR 2024 · 被引用 441 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- Rainbow Teaming: Open-Ended Generation of Diverse Adversarial PromptsMikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro 等NeurIPS 2024 · 被引用 231 次
相关 Paper
- h4rm3l: A Language for Composable Jailbreak Attack SynthesisMoussa Koulako Bala Doumbouya, Ananjan Nandi, Gabriel Poesia, Davide Ghilardi 等ICLR 2025
- SHARP: Self-adaptive Harmful Category-aware Prompt Generation for Black-box JailbreakingYingjie Xue, Xingyou Xia, Jun Zhang, Yunbo Cao 等ACL 2026
- Red Teaming LLMs via Linguistic-Aware FuzzingShuai Yuan, Nian Luo, Jingling Sun, Yihao Huang 等FSE 2026
- CoP: Agentic Red-teaming for Large Language Models using Composition of PrinciplesChen Xiong, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2025 · 被引用 13 次
- ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign UsersGuanlin Li, Kangjie Chen, Shudong Zhang, Jie Zhang 等NeurIPS 2024 · 被引用 39 次
