Towards Safe Machine Unlearning: A Paradigm that Mitigates Performance Degradation
Shanshan Ye, Jie Lu, Guangquan Zhang
Abstract
We study the machine unlearning problem which aims to remove specific training data from a pre-trained machine learning model to allow users to exercise their 'right to be forgotten' to protect user privacy. Conventional machine unlearning methods would degrade the model performance after the unlearning procedure. To mitigate the issue, they typically rely on the access to the remaining training data to fine-tune the unlearned model to mitigate the influence of unlearning. However, accessing the remaining training data may not always be practical for different reasons (e.g., data expiration policies, storage limitations, or additional privacy constraints). Machine unlearning without access to the remaining training data poses significant challenges to retaining model performance. In this paper, we study how to unlearn specific training data from a pre-trained model without accessing the remaining training data and protect model performance without dramatically changing the model's parameters. We propose a practical method called Targeted Label Noise Injection. Intuitively, our method assigns incorrect yet controllable labels to the examples that need to be forgotten and fine-tunes the pre-trained model to learn these new labels. This strategy effectively moves the to-be-forgotten examples across the decision boundary with a small impact on the model's overall performance. We theoretically prove the effectiveness of the proposed method and empirically show that it achieves state-of-the-art unlearning performance across various datasets. Relevance: Machine unlearning is critical to user modeling because users have the 'right to be forgotten'. This paper proposes a novel paradigm for machine unlearning to mitigate model performance degradation during unlearning, which is highly relevant to the track of 'User modeling, personalization and recommendation'. CCS Concepts • Networks → Network privacy and anonymity; • Security and privacy → Usability in security and privacy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5187bc17-03d9-43ea-8696-84c58fcffde6Cited by top-tier papers4
- Harnessing Hierarchical Label Distribution Variations in Test Agnostic Long-tail RecognitionZhiyong Yang, Qianqian Xu, Zitai Wang, Sicong Li et al.ICML 2024 · 21 citations
- LightFair: Towards an Efficient Alternative for Fair T2I Diffusion via Debiasing Pre-trained Text EncodersBoyu Han, Qianqian Xu, Shilong Bao, Zhiyong Yang et al.NeurIPS 2025 · 17 citations
- PARS: Partial-Label-Learning-inspired Recommender SystemsShanshan Ye, Kezhi Lu, Guangquan Zhang, Jie LuAAAI 2026 · 1 citation
- One Image is Worth a Thousand Words: A Usability Preservable Text-Image Collaborative Erasing FrameworkFeiran Li, Qianqian Xu, Shilong Bao, Zhiyong Yang et al.ICML 2025
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Anti-Backdoor Learning: Training Clean Models on Poisoned DataYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.NeurIPS 2021 · 503 citations
- Amnesiac Machine LearningLaura Graves, Vineel Nagisetty, Vijay GaneshAAAI 2021 · 416 citations
- Towards Unbounded Machine UnlearningMeghdad Kurmanji, Peter Triantafillou, Jamie Hayes, Eleni TriantafillouNeurIPS 2023 · 363 citations
Related papers
- Certified Unlearning for Neural NetworksAnastasia Koloskova, Youssef Allouah, Animesh Jha, Rachid Guerraoui et al.ICML 2025
- Remaining-data-free Machine Unlearning by Suppressing Sample ContributionXinwen Cheng, Zhehao Huang, Wenxing Zhou, Zhengbao He et al.ICLR 2026 · 11 citations
- Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine UnlearningHongsheng Hu, Shuo Wang, Tian Dong, Minhui XueS&P 2024 · 62 citations
- When Machine Unlearning Jeopardizes PrivacyMin Chen, Zhikun Zhang, Tianhao Wang, Michael Backes et al.CCS 2021 · 146 citations
- In-Context Unlearning: Language Models as Few-Shot UnlearnersMartin Pawelczyk, Seth Neel, Himabindu LakkarajuICML 2024 · 217 citations
