Improve Safety Training of Large Language Models with Safety-Critical Singular Vectors Localization
Peijian Gu, Quan Wang, Zhendong Mao
摘要
The rapid advancement of large language models (LLMs) has brought about increased concerns regarding their safety, especially as adversaries develop jailbreak techniques to bypass LLMs' safety mechanism. Although recent work on safety training with modules such as low-rank adaptation (LoRA) to resist jailbreaks shows promise, these approaches can inadvertently degrade a model's general utility. In this paper, we propose a novel plugand-play method that mitigates the impact of safety training on model utility by explicitly locating and leveraging safety-critical singular vectors, which only contribute to safety, within the model's parameter space. We quantify the safety-criticality of each singular vector as the difference of their importance for safety and utility measured by a corresponding low-rank projection. The top scored singular vectors are located as safety-critical and are used to initialize the LoRA modules within existing safety training methods in a plug-and-play manner, thereby constraining the training updates within safety-critical parameters. Additionally, we propose a dynamic rank number determination strategy to further reduce parameter overhead. Experiments on HarmBench with multiple jailbreak methods validate the effectiveness of our approach in safety training, while evaluations on several utility benchmarks demonstrate that our method successfully mitigates the adverse impact of safety training on model utility, enhancing the utility performance of the evaluated safety training baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language ModelsFanxu Meng, Zhaohui Wang, Muhan ZhangNeurIPS 2024 · 被引用 374 次
相关 Paper
- Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsBoyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie 等ICML 2024 · 被引用 215 次
- SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual ConnectionMaithili Joshi, Palash Nandi, Tanmoy ChakrabortyEMNLP 2025 · 被引用 1 次
- AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMsDebdeep Sanyal, Manodeep Ray, Murari MandalAAAI 2026 · 被引用 2 次
- JailbreakLoRA: Your Downloaded LoRA from Sharing Platforms might be UnsafeFanjunduo Wei, Zhenheng Tang, Rongfei Zeng, Tongliang Liu 等ICLR 2026
- More Thinking, Less Talking: Internalizing Deliberative Safety into LLM ParametersGuan Wang, Xuehai Tang, Biyu Zhou, Jizhong Han 等ACL 2026
