Reconstruction Attacks on Machine Unlearning: Simple Models are Vulnerable
Martin Bertran Lopez, Shuai Tang, Michael Kearns, Jamie H. Morgenstern, Aaron Roth, Steven Z. Wu
Abstract
Machine unlearning is motivated by desire for data autonomy: a person can request to have their data's influence removed from deployed models, and those models should be updated as if they were retrained without the person's data. We show that, counter-intuitively, these updates expose individuals to high-accuracy reconstruction attacks which allow the attacker to recover their data in its entirety, even when the original models are so simple that privacy risk might not otherwise have been a concern. We show how to mount a near-perfect attack on the deleted data point from linear regression models. We then generalize our attack to other loss functions and architectures, and empirically demonstrate the effectiveness of our attacks across a wide range of datasets (capturing both tabular and image data). Our work highlights that privacy risk is significant even for extremely simple model classes when individuals can request deletion of their data from the model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b19b96b1-c03d-4d40-a546-00f125ab9e54Cited by top-tier papers7
- Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLMXiaoyu Wu, Yifei Pang, Terrance Liu, Steven Z. WuNeurIPS 2025 · 8 citations
- Reminiscence Attack on Residuals: Exploiting Approximate Machine Unlearning for PrivacyYaxin Xiao, Qingqing Ye, Li Hu, Huadi Zheng et al.ICCV 2025 · 6 citations
- Dual-View Inference Attack: Machine Unlearning Amplifies Privacy ExposureLulu Xue, Shengshan Hu, Linqiang Qian, Peijin Guo et al.AAAI 2026 · 3 citations
- WARP: Weight Teleportation for Attack-Resilient Unlearning ProtocolsMohammad Mahdi Maheri, Xavier F. Cadet, Peter Chin, Hamed HaddadiICLR 2026 · 1 citation
- ReTrace: Reinforcement Learning-Guided Reconstruction Attacks on Machine UnlearningMengyao Ma, Shuofeng Liu, Minhui Xue, Surya Nepal et al.ICLR 2026
Builds on15
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
Related papers
- Hard to Forget: Poisoning Attacks on Certified Machine UnlearningNeil G. Marchant, Benjamin I. P. Rubinstein, Scott AlfeldAAAI 2022 · 95 citations
- Towards Understanding and Enhancing Robustness of Deep Learning Models against Malicious Unlearning AttacksWei Qian, Chenxu Zhao, Wei Le, Meiyi Ma et al.KDD 2023 · 38 citations
- Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine UnlearningHongsheng Hu, Shuo Wang, Tian Dong, Minhui XueS&P 2024 · 62 citations
- Unlearn and Burn: Adversarial Machine Unlearning Requests Destroy Model AccuracyYangsibo Huang, Daogao Liu, Lynn Chua, Badih Ghazi et al.ICLR 2025
- Forget Unlearning: Towards True Data-Deletion in Machine LearningRishav Chourasia, Neil ShahICML 2023 · 73 citations
