MUter: Machine Unlearning on Adversarially Trained Models
Junxu Liu, Mingsheng Xue, Jian Lou, Xiaoyu Zhang, Li Xiong, Zhan Qin
Abstract
Machine unlearning is an emerging task of removing the influence of selected training datapoints from a trained model upon data deletion requests, which echoes the widely enforced data regulations mandating the Right to be Forgotten. Many unlearning methods have been proposed recently, achieving significant efficiency gains over the naive baseline of retraining from scratch. However, existing methods focus exclusively on unlearning from standard training models and do not apply to adversarial training models (ATMs) despite their popularity as effective defenses against adversarial examples. During adversarial training, the training data are involved in not only an outer loop for minimizing the training loss, but also an inner loop for generating the adversarial perturbation. Such bi-level optimization greatly complicates the influence measure for the data to be deleted and renders the unlearning more challenging than standard model training with single-level optimization. This paper proposes a new approach called MUter for unlearning from ATMs. We derive a closed-form unlearning step underpinned by a total Hessian-related data influence measure, while existing methods can mis-capture the data influence associated with the indirect Hessian part. We further alleviate the computational cost by introducing a series of approximations and conversions to avoid the most computationally demanding parts of Hessian inversions. The efficiency and effectiveness of MUter have been validated through experiments on four datasets using both linear and neural network models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d85755e-9d81-4bf0-85a1-b5cdb4b2899bCited by top-tier papers9
- Certified Minimax Unlearning with Generalization Rates and Deletion CapacityJiaqi Liu, Jian Lou, Zhan Qin, Kui RenNeurIPS 2023 · 38 citations
- ERASER: Machine Unlearning in MLaaS via an Inference Serving-Aware ApproachYuke Hu, Jian Lou, Jiaqi Liu, Wangze Ni et al.CCS 2024 · 14 citations
- Rethinking Adversarial Robustness in the Context of the Right to be ForgottenChenxu Zhao, Wei Qian, Yangyi Li, Aobo Chen et al.ICML 2024 · 12 citations
- Ungeneralizable ExamplesJingwen Ye, Xinchao WangCVPR 2024 · 3 citations
- FALCON: Fine-grained Activation Manipulation by Contrastive Orthogonal Unalignment for Large Language ModelJinwei Hu, Zhenglin Huang, Xiangyu Yin, Wenjie Ruan et al.NeurIPS 2025 · 3 citations
Builds on27
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
- Remember What You Want to Forget: Algorithms for Machine UnlearningAyush Sekhari, Jayadev Acharya, Gautam Kamath, Ananda Theertha SureshNeurIPS 2021 · 516 citations
Related papers
- Not All Wrong is Bad: Using Adversarial Examples for UnlearningAli Ebrahimpour Boroojeny, Hari Sundaram, Varun ChandrasekaranICML 2025
- Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking DesignYuhao Sun, Yihua Zhang, Gaowen Liu, Hongtao Xie et al.ICCV 2025 · 1 citation
- Hessian-Free Online Certified UnlearningXinbao Qiao, Meng Zhang, Ming Tang, Ermin WeiICLR 2025
- The Right to be Forgotten in Federated Learning: An Efficient Realization with Rapid RetrainingYi Liu, Lei Xu, Xingliang Yuan, Cong Wang et al.INFOCOM 2022 · 189 citations
- Certified Unlearning for Neural NetworksAnastasia Koloskova, Youssef Allouah, Animesh Jha, Rachid Guerraoui et al.ICML 2025
