Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
Mickel Liu, Liwei Jiang, Yancheng Liang, Simon Du, Yejin Choi, Tim Althoff, Natasha Jaques
Abstract
Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders patch exposed vulnerabilities. This sequential setup leads to attackers overfitting obsolete exploits while defenders perpetually lag behind emerging threats. To address this, we introduce SELF-REDTEAM, the first fully online self-play multi-agent reinforcement learning (MARL) algorithm that continuously co-evolves attacker and defender for robust safety alignment. A single policy self-plays as both attacker and defender, generating adversarial prompts and defending against them, with a reward model adjudicating outcomes. Each role uses hidden chain-of-thought for strategic planning. Grounded in two-player zero-sum game theory, we establish a theoretical safety guarantee: if the game converges to Nash Equilibrium, the defender produces safe responses against any adversarial input. Empirically, SELF-REDTEAM generalizes across five models from the Llama and Qwen families, uncovering more diverse attacks (+17.80% SBERT) and improving safety of RLHF-trained models by up to 95% across 14 benchmarks. Our work motivates a shift from reactive patching to proactive co-evolution, enabling LLM safety self-improvement via online self-play MARL. Link to Code
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Absolute Zero: Reinforced Self-play Reasoning with Zero DataAndrew Zhao, Yiran Wu, Tong Wu, Quentin Xu et al.NeurIPS 2025 · 361 citations
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement LearningBo Liu, Simon Yu, Zichen Liu, Leon Guertler et al.ICLR 2026 · 88 citations
- Search Self-Play: Pushing the Frontier of Agent Capability without SupervisionHongliang Lu, Yuhang Wen, Pengyu Cheng, Ruijin Ding et al.ICLR 2026 · 35 citations
- The Alignment Waltz: Jointly Training Agents to Collaborate for SafetyJingyu Zhang, Haozhu Wang, Eric Michael Smith, Sid Wang et al.ICLR 2026 · 11 citations
- MAGIC: A Co-Evolving Attacker–Defender Adversarial Game for Robust LLM SafetyXiaoyu Wen, Zhida He, Han Qi, Ziyu Wan et al.ICML 2026 · 10 citations
Builds on34
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
Related papers
- Safety Alignment of LMs via Non-cooperative GamesAnselm Paulus, Ilia Kulikov, Brandon Amos, REMI MUNOS et al.ICML 2026 · 4 citations
- TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety AlignmentZhewen Tan, Wenhan Yu, Jianfeng Si, Tongxin Liu et al.ACL 2026 · 2 citations
- MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teamingWeiyang Guo, Jing Li, Wenya Wang, Yu Li et al.ACL 2025
- ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward SystemJiacheng Liang, Yao Ma, Tharindu Kumarage, Satyapriya Krishna et al.ACL 2026
- AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack IntegrationAndy Zhou, Kevin Wu, Francesco Pinto, Zhaorun Chen et al.NeurIPS 2025 · 46 citations
