Towards Understanding and Enhancing Robustness of Deep Learning Models against Malicious Unlearning Attacks
Wei Qian, Chenxu Zhao, Wei Le, Meiyi Ma, Mengdi Huai
Abstract
Given the availability of abundant data, deep learning models have been advanced and become ubiquitous in the past decade. In practice, due to many different reasons (e.g., privacy, usability, and fidelity), individuals also want the trained deep models to forget some specific data. Motivated by this, machine unlearning (also known as selective data forgetting) has been intensively studied, which aims at removing the influence that any particular training sample had on the trained model during the unlearning process. However, people usually employ machine unlearning methods as trusted basic tools and rarely have any doubt about their reliability. In fact, the increasingly critical role of machine unlearning makes deep learning models susceptible to the risk of being maliciously attacked. To well understand the performance of deep learning models in malicious environments, we believe that it is critical to study the robustness of deep learning models to malicious unlearning attacks, which happen during the unlearning process. To bridge this gap, in this paper, we first demonstrate that malicious unlearning attacks pose immense threats to the security of deep learning systems. Specifically, we present a broad class of malicious unlearning attacks wherein maliciously crafted unlearning requests trigger deep learning models to misbehave on target samples in a highly controllable and predictable manner. In addition, to improve the robustness of deep learning models, we also present a general defense mechanism, which aims to identify and unlearn effective malicious unlearning requests based on their gradient influence on the unlearned models. Further, theoretical analyses are conducted to analyze the proposed methods. Extensive experiments on real-world datasets validate the vulnerabilities of deep learning models to malicious unlearning attacks and the effectiveness of the introduced defense mechanism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Backdoor Attacks via Machine UnlearningZihao Liu, Tianhao Wang, Mengdi Huai, Chenglin MiaoAAAI 2024 · 46 citations
- Static and Sequential Malicious Attacks in the Context of Selective ForgettingChenxu Zhao, Wei Qian, Rex Ying, Mengdi HuaiNeurIPS 2023 · 30 citations
- UBA-Inf: Unlearning Activated Backdoor Attack with Influence-Driven CamouflageZirui Huang, Yunlong Mao, Sheng ZhongUSENIX Security 2024 · 16 citations
- Rethinking Adversarial Robustness in the Context of the Right to be ForgottenChenxu Zhao, Wei Qian, Yangyi Li, Aobo Chen et al.ICML 2024 · 12 citations
- Data Poisoning Attacks against Conformal PredictionYangyi Li, Aobo Chen, Wei Qian, Chenxu Zhao et al.ICML 2024 · 10 citations
Builds on28
- Remember What You Want to Forget: Algorithms for Machine UnlearningAyush Sekhari, Jayadev Acharya, Gautam Kamath, Ananda Theertha SureshNeurIPS 2021 · 516 citations
- Anti-Backdoor Learning: Training Clean Models on Poisoned DataYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.NeurIPS 2021 · 503 citations
- Amnesiac Machine LearningLaura Graves, Vineel Nagisetty, Vijay GaneshAAAI 2021 · 416 citations
- Enhancing Adversarial Example Transferability With an Intermediate Level AttackQian Huang, Isay Katsman, Zeqi Gu, Horace He et al.ICCV 2019 · 293 citations
- Admix: Enhancing the Transferability of Adversarial AttacksXiaosen Wang, Xuanran He, Jingdong Wang, Kun HeICCV 2021 · 282 citations
Related papers
- Reconstruction Attacks on Machine Unlearning: Simple Models are VulnerableMartin Bertran Lopez, Shuai Tang, Michael Kearns, Jamie H. Morgenstern et al.NeurIPS 2024 · 39 citations
- Unlearn and Burn: Adversarial Machine Unlearning Requests Destroy Model AccuracyYangsibo Huang, Daogao Liu, Lynn Chua, Badih Ghazi et al.ICLR 2025
- Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine UnlearningHongsheng Hu, Shuo Wang, Tian Dong, Minhui XueS&P 2024 · 62 citations
- System-Aware Unlearning Algorithms: Use Lesser, Forget FasterLinda Lu, Ayush Sekhari, Karthik SridharanICML 2025
- Machine Unlearning for Image Retrieval: A Generative Scrubbing ApproachPeng-Fei Zhang, Guangdong Bai, Zi Huang, Xin-Shun XuACM MM 2022 · 17 citations
