Machine Unlearning via Simulated Oracle Matching
Kristian Georgiev, Roy Rinberg, Sung Min Park, Shivam Garg, Andrew Ilyas, Aleksander Madry, Seth Neel
Abstract
Machine unlearning-efficiently removing the effect of a small "forget set" of training data on a pre-trained machine learning model-has recently attracted significant research interest. Despite this interest, however, recent work shows that existing machine unlearning techniques do not hold up to thorough evaluation in non-convex settings. In this work, we introduce a new machine unlearning technique that exhibits strong empirical performance even in such challenging settings. Our starting point is the perspective that the goal of unlearning is to produce a model whose outputs are statistically indistinguishable from those of a model re-trained on all but the forget set. This perspective naturally suggests a reduction from the unlearning problem to that of data attribution, where the goal is to predict the effect of changing the training set on a model's outputs. Thus motivated, we propose the following meta-algorithm, which we call Datamodel Matching (DMM): given a trained model, we (a) use data attribution to predict the output of the model if it were re-trained on all but the forget set points; then (b) fine-tune the pre-trained model to match these predicted outputs. In a simple convex setting, we show how this approach provably outperforms a variety of iterative unlearning algorithms. Empirically, we use a combination of existing evaluations and a new metric based on the KL-divergence to show that even in nonconvex settings, DMM achieves strong unlearning performance relative to existing algorithms. An added benefit of DMM is that it is a meta-algorithm, in the sense that future advances in data attribution translate directly into better unlearning algorithms, pointing to a clear direction for future progress in unlearning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Distillation Robustifies UnlearningBruce W. Lee, Addie Foote, Alex Infanger, Leni Shor et al.NeurIPS 2025 · 15 citations
- Learnability and Privacy Vulnerability are Entangled in a Few Critical WeightsXingli Fang, Jung-Eun KimICLR 2026 · 1 citation
- PeriUn: Enhancing Unlearning by Selectively Forgetting Peripheral SamplesHee Bin Yoo, Dong-Sig Han, Jaein Kim, Byoung-Tak ZhangAAAI 2026
Builds on24
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
- Amnesiac Machine LearningLaura Graves, Vineel Nagisetty, Vijay GaneshAAAI 2021 · 416 citations
Related papers
- What makes unlearning hard and what to do about itKairan Zhao, Meghdad Kurmanji, George-Octavian Barbulescu, Eleni Triantafillou et al.NeurIPS 2024 · 115 citations
- Provable unlearning in topic modeling and downstream tasksStanley Wei, Sadhika Malladi, Sanjeev Arora, Amartya SanyalICLR 2025
- De-attribute to Forget for LLM UnlearningXinyang Lu, Jiabao Pan, Rachael Hwee Ling Sim, See-Kiong Ng et al.ICML 2026
- Hessian-Free Online Certified UnlearningXinbao Qiao, Meng Zhang, Ming Tang, Ermin WeiICLR 2025
- Unlearning-Aware MinimizationHoki Kim, Keonwoo Kim, Sungwon Chae, Sangwon YoonNeurIPS 2025 · 7 citations
