Not All Wrong is Bad: Using Adversarial Examples for Unlearning
Ali Ebrahimpour Boroojeny, Hari Sundaram, Varun Chandrasekaran
Abstract
Machine unlearning, where users can request the deletion of a forget dataset, is becoming increasingly important because of numerous privacy regulations. Initial works on "exact” unlearning (e.g., retraining) incur large computational overheads. However, while computationally inexpensive, "approximate” methods have fallen short of reaching the effectiveness of exact unlearning: models produced fail to obtain comparable accuracy and prediction confidence on both the forget and test (i.e., unseen) dataset. Exploiting this observation, we propose a new unlearning method, Adversarial Machine UNlearning (AMUN), that outperforms prior state-of-the-art (SOTA) methods for image classification. AMUN lowers the confidence of the model on the forget samples by fine-tuning the model on their corresponding adversarial examples. Adversarial examples naturally belong to the distribution imposed by the model on the input space; fine-tuning the model on the adversarial examples closest to the corresponding forget samples (a) localizes the changes to the decision boundary of the model around each forget sample and (b) avoids drastic changes to the global behavior of the model, thereby preserving the model’s accuracy on test samples. Using AMUN for unlearning a random 10% of CIFAR-10 samples, we observe that even SOTA membership inference attacks cannot do better than random guessing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Approximate Machine Unlearning through Manifold Representation Forgetting Guided by Self Mode ConnectivityWeiqi Wang, Zhiyi Tian, Chenhan Zhang, Luoyu Chen et al.KDD 2026
- Unlearning Isn’t Forgetting: Revealing Hidden Leakage in Class Unlearning EvaluationsAli Ebrahimpour-Boroojeny, Yian Wang, Hari SundaramICML 2026
- Source Models Leak What They Shouldn’t: Unlearning Zero-Shot Transfer in Domain Adaptation Through Adversarial OptimizationArnav Devalapally, Poornima Jain, Kartik Srinivas, Vineeth BalasubramanianCVPR 2026
Builds on24
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
Related papers
- MUter: Machine Unlearning on Adversarially Trained ModelsJunxu Liu, Mingsheng Xue, Jian Lou, Xiaoyu Zhang et al.ICCV 2023 · 36 citations
- Reminiscence Attack on Residuals: Exploiting Approximate Machine Unlearning for PrivacyYaxin Xiao, Qingqing Ye, Li Hu, Huadi Zheng et al.ICCV 2025 · 6 citations
- Unlearning-Aware MinimizationHoki Kim, Keonwoo Kim, Sungwon Chae, Sangwon YoonNeurIPS 2025 · 7 citations
- Certified Unlearning for Neural NetworksAnastasia Koloskova, Youssef Allouah, Animesh Jha, Rachid Guerraoui et al.ICML 2025
- On the Necessity of Auditable Algorithmic Definitions for Machine UnlearningAnvith Thudi, Hengrui Jia, Ilia Shumailov, Nicolas PapernotUSENIX Security 2022
