Safe LoRA: The Silver Lining of Reducing Safety Risks when Finetuning Large Language Models
Chia-Yi Hsu, Yu-Lin Tsai, Chih-Hsun Lin, Pin-Yu Chen, Chia-Mu Yu, Chun-Ying Huang
摘要
While large language models (LLMs) such as Llama-2 or GPT-4 have shown impressive zero-shot performance, fine-tuning is still necessary to enhance their performance for customized datasets, domain-specific tasks, or other private needs. However, fine-tuning all parameters of LLMs requires significant hardware resources, which can be impractical for typical users. Therefore, parameter-efficient fine-tuning such as LoRA have emerged, allowing users to fine-tune LLMs without the need for considerable computing resources, with little performance degradation compared to fine-tuning all parameters. Unfortunately, recent studies indicate that fine-tuning can increase the risk to the safety of LLMs, even when data does not contain malicious content. To address this challenge, we propose Safe LoRA, a simple one-liner patch to the original LoRA implementation by introducing the projection of LoRA weights from selected layers to the safety-aligned subspace, effectively reducing the safety risks in LLM fine-tuning while maintaining utility. It is worth noting that Safe LoRA is a training-free and data-free approach, as it only requires the knowledge of the weights from the base and aligned LLMs. Our extensive experiments demonstrate that when fine-tuning on purely malicious data, Safe LoRA retains similar safety performance as the original aligned model. Moreover, when the fine-tuning dataset contains a mixture of both benign and malicious data, Safe LoRA mitigates the negative effect made by malicious data while preserving performance on downstream tasks. Our codes are available at https://github.com/IBM/SafeLoRA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper58
- Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning AttackTiansheng Huang, Sihao Hu, Fatih Ilhan, Selim F. Tekin 等NeurIPS 2024 · 被引用 113 次
- Representation Noising: A Defence Mechanism Against Harmful FinetuningDomenic Rosati, Jan Wehner, Kai Williams, Lukasz Bartoszcze 等NeurIPS 2024 · 被引用 107 次
- AgentAuditor: Human-level Safety and Security Evaluation for LLM AgentsHanjun Luo, Shenyu Dai, Chiming Ni, Xinfeng Li 等NeurIPS 2025 · 被引用 98 次
- Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language ModelsShengyun Peng, Pin-Yu Chen, Matthew Hull, Duen Horng ChauNeurIPS 2024 · 被引用 68 次
- Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety NeuronsJianhui Chen, Xiaozhi Wang, Zijun Yao, Yushi Bai 等NeurIPS 2025 · 被引用 53 次
它引用的顶会 Paper26
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
相关 Paper
- SaLoRA: Safety-Alignment Preserved Low-Rank AdaptationMingjie Li, Wai Man Si, Michael Backes, Yang Zhang 等ICLR 2025
- NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-TuningXin Yi, Shunfan Zheng, Linlin Wang, Gerard de Melo 等AAAI 2025 · 被引用 38 次
- Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace AdaptationDianyun Wang, Qingsen Ma, Yuhu Shang, Zhifeng Lu 等ACL 2026 · 被引用 2 次
- Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMsWeixiang Zhao, Yulin Hu, Yang Deng, Jiahe Guo 等ACL 2025 · 被引用 23 次
- Jailbreak to Protect: Buffering Harmful Fine-Tuning via Temporary Jailbreaking LoRA in Large Language ModelsSeokil Ham, Jaehyuk Jang, Wonjun Lee, Changick KimICML 2026
