Safety Alignment of LMs via Non-cooperative Games
Anselm Paulus, Ilia Kulikov, Brandon Amos, REMI MUNOS, Ivan Evtimov, Kamalika Chaudhuri, Arman Zharmagambetov
摘要
Ensuring the safety of language models (LMs) while maintaining their usefulness remains a critical challenge in AI alignment. Current approaches rely on sequential adversarial training: generating adversarial prompts and fine-tuning LMs to defend against them. We introduce a different paradigm: framing safety alignment as a non-zero-sum game between an Attacker LM and a Defender LM trained jointly via online reinforcement learning. Each LM continuously adapts to the other's evolving strategies, driving iterative improvement. Our method uses a preferencebased reward signal derived from pairwise comparisons instead of point-wise scores, providing more robust supervision and potentially reducing reward hacking. Our RL recipe, AdvGame, shifts the Pareto frontier of safety and utility, yielding a Defender LM that is simultaneously more helpful and more resilient to adversarial attacks. In addition, the resulting Attacker LM converges into a strong, general-purpose red-teaming agent that can be directly deployed to probe arbitrary target models. Code at github.com/ facebookresearch/advgame.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language ModelsMickel Liu, Liwei Jiang, Yancheng Liang, Simon Du 等ICML 2026 · 被引用 34 次
- MAGIC: A Co-Evolving Attacker–Defender Adversarial Game for Robust LLM SafetyXiaoyu Wen, Zhida He, Han Qi, Ziyu Wan 等ICML 2026 · 被引用 10 次
- Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained OptimizationTuan Nguyen, Long Tran-ThanhICML 2026
它引用的顶会 Paper18
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
相关 Paper
- TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety AlignmentZhewen Tan, Wenhan Yu, Jianfeng Si, Tongxin Liu 等ACL 2026 · 被引用 2 次
- The Alignment Waltz: Jointly Training Agents to Collaborate for SafetyJingyu Zhang, Haozhu Wang, Eric Michael Smith, Sid Wang 等ICLR 2026 · 被引用 11 次
- AGD: Adversarial Game Defense Against Jailbreak Attacks in Large Language ModelsShilong Pan, Zhiliang Tian, Zhen Huang, Wanlong Yu 等ACL 2025 · 被引用 2 次
- MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teamingWeiyang Guo, Jing Li, Wenya Wang, Yu Li 等ACL 2025
- ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward SystemJiacheng Liang, Yao Ma, Tharindu Kumarage, Satyapriya Krishna 等ACL 2026
