When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided Search
Xuan Chen, Yuzhou Nie, Wenbo Guo, Xiangyu Zhang
Abstract
Recent studies developed jailbreaking attacks, which construct jailbreaking prompts to fool LLMs into responding to harmful questions. Early-stage jailbreaking attacks require access to model internals or significant human efforts. More advanced attacks utilize genetic algorithms for automatic and black-box attacks. However, the random nature of genetic algorithms significantly limits the effectiveness of these attacks. In this paper, we propose RLbreaker, a black-box jailbreaking attack driven by deep reinforcement learning (DRL). We model jailbreaking as a search problem and design an RL agent to guide the search, which is more effective and has less randomness than stochastic search, such as genetic algorithms. Specifically, we design a customized DRL system for the jailbreaking problem, including a novel reward function and a customized proximal policy optimization (PPO) algorithm. Through extensive experiments, we demonstrate that RLbreaker is much more effective than existing jailbreaking attacks against six state-of-the-art (SOTA) LLMs. We also show that RLbreaker is robust against three SOTA defenses and its trained agents can transfer across different LLMs. We further validate the key design choices of RLbreaker via a comprehensive ablation study.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a342fcd-2064-4940-a86b-10ea34ed6c72Cited by top-tier papers24
- AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack IntegrationAndy Zhou, Kevin Wu, Francesco Pinto, Zhaorun Chen et al.NeurIPS 2025 · 46 citations
- Bileve: Securing Text Provenance in Large Language Models Against Spoofing with Bi-level SignatureTong Zhou, Xuandong Zhao, Xiaolin Xu, Shaolei RenNeurIPS 2024 · 30 citations
- Sok: Evaluating Jailbreak Guardrails for Large Language ModelsXunguang Wang, Zhenlan Ji, Wenxuan Wang, Zongjie Li et al.S&P 2026 · 27 citations
- CoP: Agentic Red-teaming for Large Language Models using Composition of PrinciplesChen Xiong, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2025 · 13 citations
- Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired SearchXun Huang, Simeng Qin, Xiaoshuang Jia, Ranjie Duan et al.ICLR 2026 · 9 citations
Builds on25
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- Robust Prompt Optimization for Defending Language Models Against Jailbreaking AttacksAndy Zhou, Bo Li, Haohan WangNeurIPS 2024 · 198 citations
- Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-to-Image Generation ModelsYingkai Dong, Xiangtao Meng, Ning Yu, Zheng Li et al.S&P 2025
- VERA: Variational Inference Framework for Jailbreaking Large Language ModelsAnamika Lochab, Lu Yan, Patrick Pynadath, Xiangyu Zhang et al.NeurIPS 2025 · 3 citations
- FlipAttack: Jailbreak LLMs via FlippingYue Liu, Xiaoxin He, Miao Xiong, Jinlan Fu et al.ICML 2025
- Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous ConstraintsJunxiao Yang, Zhexin Zhang, Shiyao Cui, Hongning Wang et al.ACL 2025 · 6 citations
