MazeBreaker: Multi-Agent Reinforcement Learning for Dynamic Jailbreaking of LLM Security Defenses
Zhihao Lin, Wei Ma, Mingyi Zhou, Yanjie Zhao, Haoyu Wang, Yang Liu, Jun Wang, Li Li
摘要
Warning:This paper contains harmful LLM responses.
In recent years, the application of Large Language Models (LLMs) has become increasingly widespread, along with growing concerns about their security. To assess the security of LLMs, researchers have proposed various jailbreak attack algorithms, but only rely on the models' internal information or face limitations in exploring the unsafe behavior, highlighting the need for a more adaptive and generalizable approach. Inspired by the game of rats escaping a maze, we introduce a novel jailbreak attack approach, MazeBreaker, where attackers dynamically learn to find the exit based on feedback and their accumulated experience to compromise the target LLMs' security defenses. Our method is the first to systematically learn from the feedback of attack attempts on target LLMs through a multi-agent reinforcement learning system, enabling strategic exploration of the model's unsafe boundaries without a reference oracle. We compared our approach with six state-of-the-art jailbreak attack methods, testing it on 13 different architectures of open-source and commercial models. The results show that our method performs exceptionally well in terms of attack effectiveness, especially for the commercial models (GPT-3.5-turbo, GPT-4o-mini, GLM-4-air and Claude-3.5-sonnet) with strong safety alignment. We hope this study will help academia and industry better test the security of large language models and promote adherence to safety and ethical standards. Code and data are available on our repository: https://anonymous.4open.science/r/MazeBreaker.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang 等ICLR 2024 · 被引用 441 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- LLM-Fuzzer: Scaling Assessment of Large Language Model JailbreaksJiahao Yu, Xingwei Lin, Zheng Yu, Xinyu XingUSENIX Security 2024 · 被引用 83 次
相关 Paper
- EquaCode: A Multi-Strategy Jailbreak Approach for Large Language Models via Equation Solving and Code CompletionZhen Liang, Hai Huang, Zhengkui ChenAAAI 2026 · 被引用 1 次
- Smoke and Mirrors: Jailbreaking LLM-based Code Generation via Implicit Malicious PromptsSheng Ouyang, Yihao Qin, Bo Lin, Liqian Chen 等ICSE 2026
- Look Before You Leap: Enhance Attention and Vigilance Regarding Harmful Content with GuidelineLLMShaoqing Zhang, Zhuosheng Zhang, Kehai Chen, Rongxiang Weng 等AAAI 2025 · 被引用 1 次
- One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMsLinbao Li, Yannan Liu, Daojing He, Yu LiICLR 2025
- MASTERKEY: Automated Jailbreaking of Large Language Model ChatbotsGelei Deng, Yi Liu, Yuekang Li, Kailong Wang 等NDSS 2024
