LSSF: Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion
Guanghao Zhou, Panjia Qiu, Cen Chen, Hongyu Li, Jason Chu, Xin Zhang, Jun Zhou
摘要
The safety mechanisms of large language models (LLMs) exhibit notable fragility, as even fine-tuning on datasets without harmful content may still undermine their safety capabilities. Meanwhile, existing safety alignment methods predominantly rely on the fine-tuning process, which inadvertently leads to the increased complexity and computational resources required. To address these issues, we introduce LSSF, a novel safety re-alignment framework with Low-Rank Safety Subspace Fusion. Our proposed method exploits the low-rank characteristics of safety information in LLMs by constructing a low-rank projection matrix to extract the principal components of safety vectors. Notably, this projection matrix represents the low-rank safety subspace of the LLMs, which we have observed to remain stable during fine-tuning process and is isolated from the model's general capabilities. These principal components are used to effectively restore safety alignment when combined with fine-tuned LLMs through linear arithmetic. Additionally, to account for the varying encoding densities of safety information across different layers of LLMs, we propose a novel metric called safety singular value entropy. This metric quantifies the encoding density and allows for the dynamic computation of the safety-critical rank for each safety vector. Extensive experiments demonstrate that our proposed post-hoc alignment method can effectively restore the safety alignment of fine-tuned models with minimal impact on their performance in downstream tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning PerturbationYibo Wang, Tiansheng Huang, Li Shen, Huanjin Yao 等NeurIPS 2025 · 被引用 22 次
- Safeguarding LLM Fine-tuning via Push-Pull Distributional AlignmentHaozhong Wang, Zhuo Li, Yibo Yang, He Zhao 等ACL 2026 · 被引用 1 次
- MMAligner: Safeguarding Multimodal Large Language Models through Representation CalibrationShenyi Zhang, Keyan Guo, Zihao Wang, Xuebin Li 等CCS 2026
- More Thinking, Less Talking: Internalizing Deliberative Safety into LLM ParametersGuan Wang, Xuehai Tang, Biyu Zhou, Jizhong Han 等ACL 2026
它引用的顶会 Paper17
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
相关 Paper
- Safety at One Shot: Patching Fine-Tuned LLMs with A Single InstanceJiawen Zhang, Lipeng He, Kejia Chen, Jian Lou 等ICLR 2026 · 被引用 10 次
- Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace AdaptationDianyun Wang, Qingsen Ma, Yuhu Shang, Zhifeng Lu 等ACL 2026 · 被引用 2 次
- Toward Safe Quantization-Aware Fine-tuning: Understanding and Mitigating Safety Alignment DegradationYuning Yang, Guowei Peng, Xiurui Xie, Minrui Jiang 等ICML 2026
- From Parameter Dynamics to Risk Scoring: Quantifying Sample-Level Safety Degradation in LLM Fine-tuningXiao Wang, Yifei Zhang, Yongkang Liu, Xiaocui Yang 等ICML 2026
- A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-SpaceBingjie Zhang, Yibo Yang, Renzhe, Dandan Guo 等ICLR 2026 · 被引用 12 次
