Backdoor Defense with Machine Unlearning
Yang Liu, Mingyuan Fan, Cen Chen, Ximeng Liu, Zhuo Ma, Li Wang, Jianfeng Ma
Abstract
Backdoor injection attack is an emerging threat to the security of neural networks, however, there still exist limited effective defense methods against the attack. In this paper, we propose BAERASER, a novel method that can erase the backdoor injected into the victim model through machine unlearning. Specifically, BAERASER mainly implements backdoor defense in two key steps. First, trigger pattern recovery is conducted to extract the trigger patterns infected by the victim model. Here, the trigger pattern recovery problem is equivalent to the one of extracting an unknown noise distribution from the victim model, which can be easily resolved by the entropy maximization based generative model. Subsequently, BAERASER leverages these recovered trigger patterns to reverse the backdoor injection procedure and induce the victim model to erase the polluted memories through a newly designed gradient ascent based machine unlearning method. Compared with the previous machine unlearning solutions, the proposed approach gets rid of the reliance on the full access to training data for retraining and shows higher effectiveness on backdoor erasing than existing fine-tuning or pruning methods. Moreover, experiments show that BAERASER can averagely lower the attack success rates of three kinds of state-of-the-art backdoor attacks by 99% on four benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e60b3a40-2f53-4d86-a2b5-4a242620a352Cited by top-tier papers42
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 365 citations
- SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and GenerationChongyu Fan, Jiancheng Liu, Yihua Zhang, Eric Wong et al.ICLR 2024 · 351 citations
- Model Sparsity Can Simplify Machine UnlearningJinghan Jia, Jiancheng Liu, Parikshit Ram, Yuguang Yao et al.NeurIPS 2023 · 293 citations
- Simplicity Prevails: Rethinking Negative Preference Optimization for LLM UnlearningChongyu Fan, Jiancheng Liu, Licong Lin, Jinghan Jia et al.NeurIPS 2025 · 182 citations
- Fast Federated Machine Unlearning with Nonlinear Functional TheoryTianshi Che, Yang Zhou, Zijie Zhang, Lingjuan Lyu et al.ICML 2023 · 77 citations
Builds on10
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.ICLR 2021 · 548 citations
- Robust anomaly detection and backdoor attack detection via differential privacyMin Du, Ruoxi Jia, Dawn SongICLR 2020 · 194 citations
Related papers
- Backdoor Attacks via Machine UnlearningZihao Liu, Tianhao Wang, Mengdi Huai, Chenglin MiaoAAAI 2024 · 46 citations
- Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine UnlearningBaogang Song, Dongdong Zhao, Jianwen Xiang, Qiben Xu et al.AAAI 2026
- Reconstructive Neuron Pruning for Backdoor DefenseYige Li, Xixiang Lyu, Xingjun Ma, Nodens Koren et al.ICML 2023 · 86 citations
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 601 citations
- LIRA: Learnable, Imperceptible and Robust Backdoor AttacksKhoa D. Doan, Yingjie Lao, Weijie Zhao, Ping LiICCV 2021 · 313 citations
