Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning Attack
Tiansheng Huang, Gautam Bhattacharya, Pratik Joshi, Joshua Kimball, Ling Liu
Abstract
Safety aligned Large Language Models (LLMs) are vulnerable to harmful fine-tuning attacks -a few harmful data mixed in the fine-tuning dataset can break the LLMs's safety alignment. While several defenses have been proposed, our evaluation shows that existing defenses fail when some specific training hyper-parameters are chosen -a large learning rate or a large number of training epochs in the fine-tuning stage can easily invalidate the defense. To this end, we propose Antidote, a post-fine-tuning stage solution, which remains agnostic to the training hyperparameters in the fine-tuning stage. Antidote relies on the philosophy that by removing the harmful parameters, the harmful model can be recovered from the harmful behaviors, regardless of how those harmful parameters are formed in the fine-tuning stage. With this philosophy, we introduce a one-shot pruning stage after harmful fine-tuning to remove the harmful weights that are responsible for the generation of harmful content. Despite its embarrassing simplicity, empirical results show that Antidote can reduce harmful score while maintaining accuracy on downstream tasks. Code is available at https: //github.com/git-disl/Antidote .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efd70f59-ce9c-46ee-9d2d-d6fa7296f66bCited by top-tier papers11
- Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuningBoheng Li, Renjie Gu, Junjie Wang, Leyi Qi et al.NeurIPS 2025 · 15 citations
- A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-SpaceBingjie Zhang, Yibo Yang, Renzhe, Dandan Guo et al.ICLR 2026 · 12 citations
- Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient InfluenceQuoc Minh Nguyen, Trung Le, Jing Wu, Anh Tuan Bui et al.ICLR 2026 · 10 citations
- Safety at One Shot: Patching Fine-Tuned LLMs with A Single InstanceJiawen Zhang, Lipeng He, Kejia Chen, Jian Lou et al.ICLR 2026 · 10 citations
- Understanding and Preserving Safety in Fine-Tuned LLMsJiawen Zhang, Yangfan Hu, Kejia Chen, Lipeng He et al.CCS 2026 · 7 citations
Builds on28
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger et al.ICLR 2024 · 373 citations
Related papers
- Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful PerturbationTiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin et al.ICLR 2025
- Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language ModelsYongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang et al.ICML 2024 · 140 citations
- SDD: Self-Degraded Defense against Malicious Fine-tuningZixuan Chen, Weikai Lu, Xin Lin, Ziqian ZengACL 2025
- AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMsDebdeep Sanyal, Manodeep Ray, Murari MandalAAAI 2026 · 2 citations
- Safety Anchor: Defending Harmful Fine-tuning via Geometric BottlenecksGuoxin Lu, Letian Sha, Qing Wang, Peijie Sun et al.ICML 2026
