Safety Misalignment Against Large Language Models
Yichen Gong, Delong Ran, Xinlei He, Tianshuo Cong, Anyu Wang, Xiaoyun Wang
摘要
—The safety alignment of Large Language Models (LLMs) is crucial to prevent unsafe content that violates human values. To ensure this, it is essential to evaluate the robustness of their alignment against diverse malicious attacks. However, the lack of a large-scale, unified measurement framework hinders a comprehensive understanding of potential vulnerabilities. To fill this gap, this paper presents the first comprehensive evaluation of existing and newly proposed safety misalignment methods for LLMs. Specifically, we investigate four research questions: (1) evaluating the robustness of LLMs with different alignment strategies, (2) identifying the most effective misalignment method, (3) determining key factors that influence misalignment effectiveness, and (4) exploring various defenses. The safety misalignment attacks in our paper include system-prompt modification, model fine-tuning, and model editing. Our findings show that Supervised Fine-Tuning is the most potent attack but requires harmful model responses. In contrast, our novel Self-Supervised Representation Attack (SSRA) achieves significant misalignment without harmful responses. We also examine defensive mechanisms such as safety data filter, model detoxification, and our proposed Self-Supervised
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMsZijian Ling, Pingyi Hu, Xiuyong Gao, Xiaojing Ma 等USENIX Security 2026 · 被引用 185 次
- Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning PerturbationYibo Wang, Tiansheng Huang, Li Shen, Huanjin Yao 等NeurIPS 2025 · 被引用 22 次
- Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-TuningWeitao Feng, Lixu Wang, Peizhuo Lv, Tianyi Wei 等CCS 2026 · 被引用 9 次
- Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention SinkGuozhi Liu, Weiwei Lin, Tiansheng Huang, Ruichao Mo 等ICML 2026 · 被引用 5 次
- Safeguarding LLM Fine-tuning via Push-Pull Distributional AlignmentHaozhong Wang, Zhuo Li, Yibo Yang, He Zhao 等ACL 2026 · 被引用 1 次
它引用的顶会 Paper36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- SDD: Self-Degraded Defense against Malicious Fine-tuningZixuan Chen, Weikai Lu, Xin Lin, Ziqian ZengACL 2025
- Safety Alignment Should be Made More Than Just a Few Tokens DeepXiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma 等ICLR 2025
- Benign Samples Matter! Fine-tuning On Outlier Benign Samples Severely Breaks SafetyZihan Guan, Mengxuan Hu, Ronghang Zhu, Sheng Li 等ICML 2025
- Safety Alignment Can Be Not Superficial With Explicit Safety SignalsJianwei Li, Jung-Eun KimICML 2025
- Defending Against Alignment-Breaking Attacks via Robustly Aligned LLMBochuan Cao, Yuanpu Cao, Lu Lin, Jinghui ChenACL 2024 · 被引用 34 次
