Learning to Unlearn: Instance-Wise Unlearning for Pre-trained Classifiers
Sungmin Cha, Sungjun Cho, Dasol Hwang, Honglak Lee, Taesup Moon, Moontae Lee
Abstract
Since the recent advent of regulations for data protection (e.g., the General Data Protection Regulation), there has been increasing demand in deleting information learned from sensitive data in pre-trained models without retraining from scratch. The inherent vulnerability of neural networks towards adversarial attacks and unfairness also calls for a robust method to remove or correct information in an instance-wise fashion, while retaining the predictive performance across remaining data. To this end, we consider instance-wise unlearning, of which the goal is to delete information on a set of instances from a pre-trained model, by either misclassifying each instance away from its original prediction or relabeling the instance to a different label. We also propose two methods that reduce forgetting on the remaining data: 1) utilizing adversarial examples to overcome forgetting at the representation-level and 2) leveraging weight importance metrics to pinpoint network parameters guilty of propagating unwanted information. Both methods only require the pre-trained model and data instances to forget, allowing painless application to real-life settings where the entire training set is unavailable. Through extensive experimentation on various image classification benchmarks, we show that our approach effectively preserves knowledge of remaining data while unlearning given instances in both single-task and continual unlearning scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4faaa1b2-6b69-4229-840e-cee091d3b838Cited by top-tier papers36
- Single Image Unlearning: Efficient Machine Unlearning in Multimodal Large Language ModelsJiaqi Li, Qianshan Wei, Chuanyi Zhang, Guilin Qi et al.NeurIPS 2024 · 62 citations
- On Effects of Steering Latent Representation for Large Language Model UnlearningHuu-Tien Dang, Tin Pham, Hoang Thanh-Tung, Naoya InoueAAAI 2025 · 33 citations
- Towards Safe Machine Unlearning: A Paradigm that Mitigates Performance DegradationShanshan Ye, Jie Lu, Guangquan ZhangWWW 2025 · 13 citations
- Distribution-Level Feature Distancing for Machine Unlearning: Towards a Better Trade-off Between Model Utility and ForgettingDasol Choi, Dongbin NaAAAI 2025 · 9 citations
- Label Smoothing Improves Machine UnlearningZonglin Di, Zhaowei Zhu, Jinghan Jia, Jiancheng Liu et al.ICLR 2026 · 8 citations
Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Amnesiac Machine LearningLaura Graves, Vineel Nagisetty, Vijay GaneshAAAI 2021 · 416 citations
- Fairness without Demographics through Adversarially Reweighted LearningPreethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee et al.NeurIPS 2020 · 406 citations
Related papers
- Partially Blinded Unlearning: Class Unlearning for Deep Networks from Bayesian PerspectiveSubhodip Panda, Shashwat Sourav, Prathosh A. P.AAAI 2025 · 3 citations
- Is Graph Unlearning Ready for Practice? A Benchmark on Efficiency, Utility, and ForgettingSamyak Jain, Ronak Kalvani, sainyam galhotra, Sayan RanuICLR 2026
- In-Context Unlearning: Language Models as Few-Shot UnlearnersMartin Pawelczyk, Seth Neel, Himabindu LakkarajuICML 2024 · 217 citations
- Decoupled Distillation to Erase: A General Unlearning Method for Any Class-centric TasksYu Zhou, Dian Zheng, Qijie Mo, Renjie Lu et al.CVPR 2025
- Remaining-data-free Machine Unlearning by Suppressing Sample ContributionXinwen Cheng, Zhehao Huang, Wenxing Zhou, Zhengbao He et al.ICLR 2026 · 11 citations
