SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models
Muxi Diao, Rumei Li, Shiyang Liu, Guogang Liao, Jingang Wang, Xunliang Cai, Weiran Xu
摘要
As Large Language Models (LLMs) continue to advance in capability and influence, ensuring their safety and preventing harmful outputs has become crucial. A promising approach to address these concerns involves training models to automatically generate adversarial prompts for red teaming. However, the evolving subtlety of vulnerabilities in LLMs challenges the effectiveness of current adversarial methods, which struggle to generate diverse, complex prompts and dynamically explore the weaknesses of these models. To tackle these challenges, we introduce the Self-Evolving Adversarial Safety (SEAS) optimization framework, which includes both a SEAS dataset and a SEAS pipeline. The SEAS dataset comprises complex adversarial prompts, while the SEAS pipeline operates through three stages: Initialization, Attack, and Adversarial Optimization. This framework generates a diverse range of adversarial prompts and dynamically explores the model's vulnerabilities to enhance its safety. Our contributions include a novel adversarial framework, a comprehensive safety dataset, and empirical evidence demonstrating the effectiveness of SEAS. After three iterations, the model achieves a safety level comparable to GPT-4. Our code and datasets are released at https://SEAS-LLM.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level PrecisionZhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing 等AAAI 2026 · 被引用 4 次
- MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language ModelsSiyu Yan, Long Zeng, Xuecheng Wu, Chengcheng Han 等EMNLP 2025 · 被引用 1 次
- CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science MasteryXiaoshuai Song, Muxi Diao, Guanting Dong, Zhengyang Wang 等ICLR 2025
- Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety AlignmentJiajia Li, Xiaoyu Wen, Shuyue Hu, Qiaosheng Zhang 等ICML 2026
- Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMsXiang Zheng, YUTAO WU, Hanxun Huang, Yige Li 等ICML 2026
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt GenerationHuawei Zheng, Xinqi Jiang, Sen Yang, Shouling Ji 等ACL 2026 · 被引用 1 次
- Automated Red Teaming with GOAT: the Generative Offensive Agent TesterMaya Pavlova, Erik Brinkman, Krithika Iyer, Vítor Albiero 等ICML 2025
- Automated Red Teaming for Text-to-Image Models Through Feedback-Guided Prompt Iteration with Vision-Language ModelsWei Xu, Kangjie Chen, Jiawei Qiu, Yuyang Zhang 等ICCV 2025 · 被引用 3 次
- Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMsZhiyang Chen, Tara Saba, Xun Deng, Xujie Si 等ICML 2026
- ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign UsersGuanlin Li, Kangjie Chen, Shudong Zhang, Jie Zhang 等NeurIPS 2024 · 被引用 39 次
