Backdoor Attacks via Machine Unlearning
Zihao Liu, Tianhao Wang, Mengdi Huai, Chenglin Miao
Abstract
As a new paradigm to erase data from a model and protect user privacy, machine unlearning has drawn significant attention. However, existing studies on machine unlearning mainly focus on its effectiveness and efficiency, neglecting the security challenges introduced by this technique. In this paper, we aim to bridge this gap and study the possibility of conducting malicious attacks leveraging machine unlearning. Specifically, we consider the backdoor attack via machine unlearning, where an attacker seeks to inject a backdoor in the unlearned model by submitting malicious unlearning requests, so that the prediction made by the unlearned model can be changed when a particular trigger presents. In our study, we propose two attack approaches. The first attack approach does not require the attacker to poison any training data of the model. The attacker can achieve the attack goal only by requesting to unlearn a small subset of his contributed training data. The second approach allows the attacker to poison a few training instances with a pre-defined trigger upfront, and then activate the attack via submitting a malicious unlearning request. Both attack approaches are proposed with the goal of maximizing the attack utility while ensuring attack stealthiness. The effectiveness of the proposed attacks is demonstrated with different machine unlearning algorithms as well as different models on different datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e20b4f2b-a41c-489f-917c-984f6fabe011Cited by top-tier papers12
- Backdoor Attacks Against No-Reference Image Quality Assessment Models via a Scalable TriggerYi Yu, Song Xia, Xun Lin, Wenhan Yang et al.AAAI 2025 · 15 citations
- Label-Free Backdoor Attacks in Vertical Federated LearningWei Shen, Wenke Huang, Guancheng Wan, Mang YeAAAI 2025 · 15 citations
- Rethinking Adversarial Robustness in the Context of the Right to be ForgottenChenxu Zhao, Wei Qian, Yangyi Li, Aobo Chen et al.ICML 2024 · 12 citations
- Data Poisoning Attacks against Conformal PredictionYangyi Li, Aobo Chen, Wei Qian, Chenxu Zhao et al.ICML 2024 · 10 citations
- StealthAttack: Robust 3D Gaussian Splatting Poisoning via Density-Guided IllusionsBo-Hsu Ke, You-Zhe Xie, Yu-Lun Liu, Wei-Chen ChiuICCV 2025 · 3 citations
Builds on13
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
- Anti-Backdoor Learning: Training Clean Models on Poisoned DataYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.NeurIPS 2021 · 503 citations
- Machine Unlearning for Random ForestsJonathan Brophy, Daniel LowdICML 2021 · 222 citations
- Composite Backdoor Attack for Deep Neural Network by Mixing Existing Benign FeaturesJunyu Lin, Lei Xu, Yingqi Liu, Xiangyu ZhangCCS 2020 · 197 citations
Related papers
- UBA-Inf: Unlearning Activated Backdoor Attack with Influence-Driven CamouflageZirui Huang, Yunlong Mao, Sheng ZhongUSENIX Security 2024 · 16 citations
- Backdoor Defense with Machine UnlearningYang Liu, Mingyuan Fan, Cen Chen, Ximeng Liu et al.INFOCOM 2022 · 89 citations
- Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine UnlearningBaogang Song, Dongdong Zhao, Jianwen Xiang, Qiben Xu et al.AAAI 2026
- Shared Adversarial Unlearning: Backdoor Mitigation by Unlearning Shared Adversarial ExamplesShaokui Wei, Mingda Zhang, Hongyuan Zha, Baoyuan WuNeurIPS 2023 · 69 citations
- Hard to Forget: Poisoning Attacks on Certified Machine UnlearningNeil G. Marchant, Benjamin I. P. Rubinstein, Scott AlfeldAAAI 2022 · 95 citations
