MazeBreaker: Multi-Agent Reinforcement Learning for Dynamic Jailbreaking of LLM Security Defenses
Zhihao Lin, Wei Ma, Mingyi Zhou, Yanjie Zhao, Haoyu Wang, Yang Liu, Jun Wang, Li Li
Abstract
Warning:This paper contains harmful LLM responses.
In recent years, the application of Large Language Models (LLMs) has become increasingly widespread, along with growing concerns about their security. To assess the security of LLMs, researchers have proposed various jailbreak attack algorithms, but only rely on the models' internal information or face limitations in exploring the unsafe behavior, highlighting the need for a more adaptive and generalizable approach. Inspired by the game of rats escaping a maze, we introduce a novel jailbreak attack approach, MazeBreaker, where attackers dynamically learn to find the exit based on feedback and their accumulated experience to compromise the target LLMs' security defenses. Our method is the first to systematically learn from the feedback of attack attempts on target LLMs through a multi-agent reinforcement learning system, enabling strategic exploration of the model's unsafe boundaries without a reference oracle. We compared our approach with six state-of-the-art jailbreak attack methods, testing it on 13 different architectures of open-source and commercial models. The results show that our method performs exceptionally well in terms of attack effectiveness, especially for the commercial models (GPT-3.5-turbo, GPT-4o-mini, GLM-4-air and Claude-3.5-sonnet) with strong safety alignment. We hope this study will help academia and industry better test the security of large language models and promote adherence to safety and ethical standards. Code and data are available on our repository: https://anonymous.4open.science/r/MazeBreaker.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang et al.ICLR 2024 · 441 citations
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
- LLM-Fuzzer: Scaling Assessment of Large Language Model JailbreaksJiahao Yu, Xingwei Lin, Zheng Yu, Xinyu XingUSENIX Security 2024 · 83 citations
Related papers
- EquaCode: A Multi-Strategy Jailbreak Approach for Large Language Models via Equation Solving and Code CompletionZhen Liang, Hai Huang, Zhengkui ChenAAAI 2026 · 1 citation
- Smoke and Mirrors: Jailbreaking LLM-based Code Generation via Implicit Malicious PromptsSheng Ouyang, Yihao Qin, Bo Lin, Liqian Chen et al.ICSE 2026
- Look Before You Leap: Enhance Attention and Vigilance Regarding Harmful Content with GuidelineLLMShaoqing Zhang, Zhuosheng Zhang, Kehai Chen, Rongxiang Weng et al.AAAI 2025 · 1 citation
- One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMsLinbao Li, Yannan Liu, Daojing He, Yu LiICLR 2025
- MASTERKEY: Automated Jailbreaking of Large Language Model ChatbotsGelei Deng, Yi Liu, Yuekang Li, Kailong Wang et al.NDSS 2024
