Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
Danny Halawi, Alexander Wei, Eric Wallace, Tony Tong Wang, Nika Haghtalab, Jacob Steinhardt
Abstract
Black-box finetuning is an emerging interface for adapting state-of-the-art language models to user needs. However, such access may also let malicious actors undermine model safety. To demonstrate the challenge of defending finetuning interfaces, we introduce covert malicious finetuning, a method to compromise model safety via finetuning while evading detection. Our method constructs a malicious dataset where every individual datapoint appears innocuous, but finetuning on the dataset teaches the model to respond to encoded harmful requests with encoded harmful responses. Applied to GPT-4, our method produces a finetuned model that acts on harmful instructions 99% of the time and avoids detection by defense mechanisms such as dataset inspection, safety evaluations, and input/output classifiers. Our findings question whether black-box finetuning access can be secured against sophisticated adversaries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6db5d12d-9530-4617-8e19-892dc980aa01Cited by top-tier papers41
- Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning AttackTiansheng Huang, Sihao Hu, Fatih Ilhan, Selim F. Tekin et al.NeurIPS 2024 · 113 citations
- Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal TrainingYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang et al.ACL 2025 · 65 citations
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMsKyle O'Brien, Stephen Casper, Quentin Anthony, Tomek Korbak et al.ICLR 2026 · 59 citations
- Estimating Worst-Case Frontier Risks of Open-Weight LLMsEric Wallace, Olivia Watkins, Miles Wang, Kai Chen et al.ICLR 2026 · 31 citations
- Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning DatasetsLei Hsiung, Tianyu Pang, Yung-Chen Tang, Linyue Song et al.ACL 2026 · 22 citations
Builds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang et al.ICLR 2024 · 441 citations
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger et al.ICLR 2024 · 373 citations
Related papers
- Invisible Safety Threat: Malicious Finetuning for LLM via SteganographyGuangnian Wan, Xinyin Ma, Gongfan Fang, Xinchao WangICLR 2026 · 4 citations
- No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety MechanismsJoshua Kazdan, Abhay Puri, Rylan Schaeffer, Lisa Yu et al.ICLR 2026
- Fundamental Limitations in Pointwise Defences of LLM Finetuning APIsXander Davies, Eric Winsor, Alexandra Souly, Tomek Korbak et al.NeurIPS 2025 · 13 citations
- Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!Zhexin Zhang, Yuhao Sun, Junxiao Yang, Shiyao Cui et al.ICLR 2026 · 5 citations
- Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language ModelsYongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang et al.ICML 2024 · 140 citations
