Machine Unlearning Fails to Remove Data Poisoning Attacks
Martin Pawelczyk, Jimmy Z. Di, Yiwei Lu, Gautam Kamath, Ayush Sekhari, Seth Neel
Abstract
We revisit the efficacy of several practical methods for approximate machine unlearning developed for large-scale deep learning. In addition to complying with data deletion requests, one often-cited potential application for unlearning methods is to remove the effects of training on poisoned data. We experimentally demonstrate that, while existing unlearning methods have been demonstrated to be effective in a number of evaluation settings (e.g., alleviating membership inference attacks), they fail to remove the effects of data poisoning, across a variety of types of poisoning attacks (indiscriminate, targeted, and a newly-introduced Gaussian poisoning attack) and models (image classifiers and LLMs); even when granted a relatively large compute budget. In order to precisely characterize unlearning efficacy, we introduce new evaluation metrics for unlearning based on data poisoning. Our results suggest that a broader perspective, including a wider variety of evaluations, are required to avoid a false sense of confidence in machine unlearning procedures for deep learning without provable guarantees. Moreover, while unlearning methods show some signs of being useful to efficiently remove poisoned datapoints without having to retrain, our work suggests that these methods are not yet "ready for prime time," and currently provide limited benefit over retraining. ⋆ MP and JZD have equal first-author contributions, and GK, AS, and SN have equal advisory contributions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6389fed-3ba1-407e-8e7c-ec203b4efd41Cited by top-tier papers18
- Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence AwarenessRongzhe Wei, Peizhi Niu, Hans Hao-Hsun Hsu, Ruihan Wu et al.NeurIPS 2025 · 20 citations
- Ascent Fails to ForgetIoannis Mavrothalassitis, Pol Puigdemont, Noam Itzhak Levi, Volkan CevherNeurIPS 2025 · 12 citations
- Train Once, Answer All: Many Pretraining Experiments for the Cost of OneSebastian Bordt, Martin PawelczykICLR 2026 · 6 citations
- Gaussian certified unlearning in high dimensions: A hypothesis testing approachAaradhya Pandey, Arnab Auddy, Haolin Zou, Arian Maleki et al.ICLR 2026 · 5 citations
- The Unseen Threat: Residual Knowledge in Machine Unlearning under Perturbed SamplesHsiang Hsu, Pradeep Niroula, Zichang He, Ivan Brugere et al.NeurIPS 2025 · 5 citations
Builds on32
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
Related papers
- Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Unlearning CompletenessCheng-Long Wang, Qi Li, Zihang Xiang, Yinzhi Cao et al.USENIX Security 2025
- A Reliable Cryptographic Framework for Empirical Machine Unlearning EvaluationYiwen Tu, Pingbang Hu, Jiaqi MaNeurIPS 2025 · 6 citations
- Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning AttacksJimmy Z. Di, Jack Douglas, Jayadev Acharya, Gautam Kamath et al.NeurIPS 2023 · 70 citations
- Towards Unbounded Machine UnlearningMeghdad Kurmanji, Peter Triantafillou, Jamie Hayes, Eleni TriantafillouNeurIPS 2023 · 363 citations
- What makes unlearning hard and what to do about itKairan Zhao, Meghdad Kurmanji, George-Octavian Barbulescu, Eleni Triantafillou et al.NeurIPS 2024 · 115 citations
