MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
Weiyang Guo, Jing Li, Wenya Wang, Yu Li, Daojing He, Jun Yu, Min Zhang
摘要
The proliferation of jailbreak attacks against large language models (LLMs) highlights the need for robust security measures. However, in multi-round dialogues, malicious intentions may be hidden in interactions, leading LLMs to be more prone to produce harmful responses. In this paper, we propose the Multi-Turn Safety Alignment (MTSA) framework, to address the challenge of securing LLMs in multi-round interactions. It consists of two stages: In the thought-guided attack learning stage, the redteam model learns about thought-guided multiround jailbreak attacks to generate adversarial prompts. In the adversarial iterative optimization stage, the red-team model and the target model continuously improve their respective capabilities in interaction. Furthermore, we introduce a multi-turn reinforcement learning algorithm based on future rewards to enhance the robustness of safety alignment. Experimental results show that the red-team model exhibits state-of-the-art attack capabilities, while the target model significantly improves its performance on safety benchmarks. Code is available at https://github.com/yuki-younai/MTSA WARNING: This paper contains potentially offensive and harmful text.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming AttacksRuohao Guo, Afshin Oroojlooyjadid, Roshan Sridhar, Miguel Ballesteros 等ICLR 2026 · 被引用 12 次
- TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process RewardsXiqiao Xiong, Ouxiang Li, Zhuo Liu, Moxin Li 等ACL 2026 · 被引用 7 次
- Safety Alignment of Large Language Models via Contrasting Safe and Harmful DistributionsXiaoyun Zhang, Zhengyue Zhao, Wenxuan Shi, Kaidi Xu 等AAAI 2026 · 被引用 4 次
- Evaluating Temporal Consistency in Multi-Turn Language ModelsYash Kumar Atri, Steven L. Johnson, Thomas HartvigsenACL 2026 · 被引用 1 次
- Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy OptimizationHuilin Zhou, Jian Zhao, Yilu Zhong, Zhen Liang 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper16
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji 等ICLR 2024 · 被引用 656 次
相关 Paper
- Multi-Turn Jailbreaking Large Language Models via Attention ShiftingXiaohu Du, Fan Mo, Ming Wen, Tu Gu 等AAAI 2025 · 被引用 26 次
- MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language ModelsSiyu Yan, Long Zeng, Xuecheng Wu, Chengcheng Han 等EMNLP 2025 · 被引用 1 次
- SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak AttacksMingqian Feng, Xiaodong Liu, Weiwei Yang, Jialin Song 等ICLR 2026 · 被引用 13 次
- DAMON: A Dialogue-Aware MCTS Framework for Jailbreaking Large Language ModelsXu Zhang, Xunjian Yin, Dinghao Jing, Huixuan Zhang 等EMNLP 2025 · 被引用 2 次
- Analogy-based Multi-Turn Jailbreak against Large Language ModelsMengjie Wu, Yihao Huang, Zhenjun Lin, Kangjie Chen 等NeurIPS 2025 · 被引用 9 次
