Lifelong Safety Alignment for Language Models
Haoyu Wang, Yifei Zhao, Zeyu Qin, Chao Du, Min Lin, Xueqian Wang, Tianyu Pang
Abstract
LLMs have made impressive progress, but their growing capabilities also expose them to highly flexible jailbreaking attacks designed to bypass safety alignment. While many existing defenses focus on known types of attacks, it is more critical to prepare LLMs for unseen attacks that may arise during deployment. To address this, we propose a lifelong safety alignment framework that enables LLMs to continuously adapt to new and evolving jailbreaking strategies. Our framework introduces a competitive setup between two components: a Meta-Attacker, trained to actively discover novel jailbreaking strategies, and a Defender, trained to resist them. To effectively warm up the Meta-Attacker, we first leverage the GPT-4o API to extract key insights from a large collection of jailbreak-related research papers. Through iterative training, the first iteration Meta-Attacker achieves a 73% attack success rate (ASR) on RR [80] and a 57% transfer ASR on LAT [53] using only single-turn attacks. Meanwhile, the Defender progressively improves its robustness and ultimately reduces the Meta-Attacker's success rate to just 7%, enabling safer and more reliable deployment of LLMs in open-ended environments. The code is available at https://github.com/sail-sg/LifelongSafetyAlignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2823e535-0bef-4fdf-8aa9-ec983eb851bbCited by top-tier papers5
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image GenerationYifu Luo, Xinhao Hu, Keyu Fan, Haoyuan Sun et al.NeurIPS 2025 · 12 citations
- Principled RL for Flow Matching Emerges from the Chunk-level Policy OptimizationYifu Luo, Haoyuan Sun, Xinhao Hu, Penghui Du et al.ICML 2026 · 12 citations
- Safety Alignment of LMs via Non-cooperative GamesAnselm Paulus, Ilia Kulikov, Brandon Amos, REMI MUNOS et al.ICML 2026 · 4 citations
- Safety Reasoning with GuidelinesHaoyu Wang, Zeyu Qin, Li Shen, Xueqian Wang et al.ICML 2025
- Reflector: Internalizing Step-wise Reflection against Indirect JailbreaksJiachen Ma, Jiawen Zhang, Xiangtian Li, Bo Zou et al.ICML 2026
Builds on37
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
Related papers
- SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical MannerXunguang Wang, Daoyuan Wu, Zhenlan Ji, Zongjie Li et al.USENIX Security 2025
- One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMsLinbao Li, Yannan Liu, Daojing He, Yu LiICLR 2025
- Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language ModelsGuangyu Yang, Jinghong Chen, Jingbiao Mei, Weizhe Lin et al.ACL 2026 · 1 citation
- Defending Large Language Models Against Jailbreaking Attacks Through Goal PrioritizationZhexin Zhang, Junxiao Yang, Pei Ke, Fei Mi et al.ACL 2024
- Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time AlignmentSoumya Suvra Ghosal, Souradip Chakraborty, Vaibhav Singh, Tianrui Guan et al.CVPR 2025
