Broadening the Backdoor Basin: Understanding LLM Backdoors Collapse and Making Backdoors Persistent
Xingyi Zhao, Tian Xie, Xiaojun Qi, Depeng Xu, Shuhan Yuan
Abstract
Large Language Models (LLMs) are vulnerable to backdoor attacks, yet we observe that many LLM backdoors do not survive when end users perform supervised fine-tuning (SFT). In this work, we provide a geometric explanation: by probing the backdoor objective under controlled weight perturbations, we find that conventional poisoning often drives the backdoor loss to a narrow and sharp basin; consequently, even modest parameter drift induced by downstream SFT can push the model out of the low-loss and high-ASR region, leading to rapid backdoor forgetting. Motivated by this insight, we propose BAD-BOOM, a resilient backdoor attack via broader smoothness minimization, which explicitly broadens and smooths the backdoor basin. BAD-BOOM extends sharpness-aware minimization with a Fisher-induced ellipsoidal constraint that allocates larger perturbation budgets to backdoor-sensitive parameters, encouraging solutions whose neighborhoods also maintain low backdoor loss. Across two threat settings, three attack scenarios, three open-source LLMs, and three trigger-free downstream SFT tasks, BAD-BOOM consistently preserves high ASR while maintaining competitive utility. The code is available at https://github.com/xingyizhao/BAD-BOOM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 665e4c01-f8d2-498e-bf20-bbbb02a75a47Builds on19
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 319 citations
- RL's Razor: Why Online Reinforcement Learning Forgets LessIdan Shenfeld, Jyothish Pari, Pulkit AgrawalICLR 2026 · 176 citations
- On the Exploitability of Instruction TuningManli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping et al.NeurIPS 2023 · 166 citations
Related papers
- BadAgent: Inserting and Activating Backdoor Attacks in LLM AgentsYifei Wang, Dizhan Xue, Shengjie Zhang, Shengsheng QianACL 2024
- Persistent Backdoor Attacks Under Continual Fine-Tuning of LLMsJing Cui, Yufei Han, Jianbin Jiao, Junge ZhangAAAI 2026
- BadEdit: Backdooring Large Language Models by Model EditingYanzhou Li, Tianlin Li, Kangjie Chen, Jian Zhang et al.ICLR 2024 · 116 citations
- Dormant Backdoor: Weaponizing Model Finetuning for Feasible Backdoor Attacks Against Pretrained ModelsRuitao Li, Jiakai Wang, Hairong Chen, Huihu Ding et al.AAAI 2026
- RepGuard: Adaptive Feature Decoupling for Robust Backdoor Defense in Large Language ModelsChenxu Niu, Jie M. Zhang, Yanbing Liu, Yunpeng Li et al.NeurIPS 2025 · 1 citation
