Backdoor Attacks via Machine Unlearning
Zihao Liu, Tianhao Wang, Mengdi Huai, Chenglin Miao
摘要
As a new paradigm to erase data from a model and protect user privacy, machine unlearning has drawn significant attention. However, existing studies on machine unlearning mainly focus on its effectiveness and efficiency, neglecting the security challenges introduced by this technique. In this paper, we aim to bridge this gap and study the possibility of conducting malicious attacks leveraging machine unlearning. Specifically, we consider the backdoor attack via machine unlearning, where an attacker seeks to inject a backdoor in the unlearned model by submitting malicious unlearning requests, so that the prediction made by the unlearned model can be changed when a particular trigger presents. In our study, we propose two attack approaches. The first attack approach does not require the attacker to poison any training data of the model. The attacker can achieve the attack goal only by requesting to unlearn a small subset of his contributed training data. The second approach allows the attacker to poison a few training instances with a pre-defined trigger upfront, and then activate the attack via submitting a malicious unlearning request. Both attack approaches are proposed with the goal of maximizing the attack utility while ensuring attack stealthiness. The effectiveness of the proposed attacks is demonstrated with different machine unlearning algorithms as well as different models on different datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Backdoor Attacks Against No-Reference Image Quality Assessment Models via a Scalable TriggerYi Yu, Song Xia, Xun Lin, Wenhan Yang 等AAAI 2025 · 被引用 15 次
- Label-Free Backdoor Attacks in Vertical Federated LearningWei Shen, Wenke Huang, Guancheng Wan, Mang YeAAAI 2025 · 被引用 15 次
- Rethinking Adversarial Robustness in the Context of the Right to be ForgottenChenxu Zhao, Wei Qian, Yangyi Li, Aobo Chen 等ICML 2024 · 被引用 12 次
- Data Poisoning Attacks against Conformal PredictionYangyi Li, Aobo Chen, Wei Qian, Chenxu Zhao 等ICML 2024 · 被引用 10 次
- StealthAttack: Robust 3D Gaussian Splatting Poisoning via Density-Guided IllusionsBo-Hsu Ke, You-Zhe Xie, Yu-Lun Liu, Wei-Chen ChiuICCV 2025 · 被引用 3 次
它引用的顶会 Paper13
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia 等S&P 2021 · 被引用 1,381 次
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 被引用 633 次
- Anti-Backdoor Learning: Training Clean Models on Poisoned DataYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu 等NeurIPS 2021 · 被引用 503 次
- Machine Unlearning for Random ForestsJonathan Brophy, Daniel LowdICML 2021 · 被引用 222 次
- Composite Backdoor Attack for Deep Neural Network by Mixing Existing Benign FeaturesJunyu Lin, Lei Xu, Yingqi Liu, Xiangyu ZhangCCS 2020 · 被引用 197 次
相关 Paper
- UBA-Inf: Unlearning Activated Backdoor Attack with Influence-Driven CamouflageZirui Huang, Yunlong Mao, Sheng ZhongUSENIX Security 2024 · 被引用 16 次
- Backdoor Defense with Machine UnlearningYang Liu, Mingyuan Fan, Cen Chen, Ximeng Liu 等INFOCOM 2022 · 被引用 89 次
- Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine UnlearningBaogang Song, Dongdong Zhao, Jianwen Xiang, Qiben Xu 等AAAI 2026
- Shared Adversarial Unlearning: Backdoor Mitigation by Unlearning Shared Adversarial ExamplesShaokui Wei, Mingda Zhang, Hongyuan Zha, Baoyuan WuNeurIPS 2023 · 被引用 69 次
- Hard to Forget: Poisoning Attacks on Certified Machine UnlearningNeil G. Marchant, Benjamin I. P. Rubinstein, Scott AlfeldAAAI 2022 · 被引用 95 次
