Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation
Yibo Wang, Tiansheng Huang, Li Shen, Huanjin Yao, Haotian Luo, Rui Liu, Naiqiang Tan, Jiaxing Huang, Dacheng Tao
Abstract
Harmful fine-tuning attack introduces significant security risks to the fine-tuning services. Main-stream defenses aim to vaccinate the model such that the later harmful fine-tuning attack is less effective. However, our evaluation results show that such defenses are fragile-with a few fine-tuning steps, the model still can learn the harmful knowledge. To this end, we do further experiment and find that an embarrassingly simple solution-adding purely random perturbations to the fine-tuned model, can recover the model from harmful behaviors, though it leads to a degradation in the model's fine-tuning performance. To address the degradation of fine-tuning performance, we further propose Panacea, which optimizes an adaptive perturbation that will be applied to the model after fine-tuning. Panacea maintains model's safety alignment performance without compromising downstream fine-tuning performance. Comprehensive experiments are conducted on different harmful ratios, fine-tuning tasks and mainstream LLMs, where the average harmful scores are reduced by up-to 21.2%, while maintaining fine-tuning performance. As a by-product, we analyze the adaptive perturbation and show that different layers in various LLMs have distinct safety affinity, which coincide with finding from several previous study. Source code available at https://github. com/w-yibo/Panacea . Vaccine [8] and RepNoise [9] are two representative defenses against the harmful fine-tuning attack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d31cd390-e129-4b44-9ab8-42509c756a62Cited by top-tier papers11
- Backdoor Cleaning without External Guidance in MLLM Fine-tuningXuankun Rong, Wenke Huang, Jian Liang, Jinhe Bi et al.NeurIPS 2025 · 39 citations
- Shape it Up! Restoring LLM Safety during FinetuningShengyun Peng, Pin-Yu Chen, Jianfeng Chi, Seongmin Lee et al.NeurIPS 2025 · 17 citations
- Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data SchedulerZixuan Hu, Li Shen, Zhenyi Wang, Yongxian Wei et al.NeurIPS 2025 · 16 citations
- Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuningBoheng Li, Renjie Gu, Junjie Wang, Leyi Qi et al.NeurIPS 2025 · 15 citations
- Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient InfluenceQuoc Minh Nguyen, Trung Le, Jing Wu, Anh Tuan Bui et al.ICLR 2026 · 10 citations
Builds on57
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger et al.ICLR 2024 · 373 citations
- Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsBoyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie et al.ICML 2024 · 215 citations
- Safe LoRA: The Silver Lining of Reducing Safety Risks when Finetuning Large Language ModelsChia-Yi Hsu, Yu-Lin Tsai, Chih-Hsun Lin, Pin-Yu Chen et al.NeurIPS 2024 · 165 citations
Related papers
- Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning AttackTiansheng Huang, Gautam Bhattacharya, Pratik Joshi, Joshua Kimball et al.ICML 2025
- Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful PerturbationTiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin et al.ICLR 2025
- OASIS: Mitigating Harmful Fine-tuning Attacks on LLMs via Orthogonal and Adaptive Safety Alignment StrategyJiayu Tang, Guowei Peng, Qiuhao Xie, Yuning Yang et al.ACL 2026
- Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning AttackTiansheng Huang, Sihao Hu, Ling LiuNeurIPS 2024 · 3 citations
- CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-TuningBiao Yi, Tiansheng Huang, Baolei Zhang, Tong Li et al.ACL 2026 · 15 citations
