Repairing Neural Networks by Leaving the Right Past Behind
Ryutaro Tanno, Melanie F. Pradier, Aditya V. Nori, Yingzhen Li
Abstract
Prediction failures of machine learning models often arise from deficiencies in training data, such as incorrect labels, outliers, and selection biases. However, such data points that are responsible for a given failure mode are generally not known a priori, let alone a mechanism for repairing the failure. This work draws on the Bayesian view of continual learning, and develops a generic framework for both, identifying training examples which have given rise to the target failure, and fixing the model through erasing information about them. This framework naturally allows leveraging recent advances in continual learning to this new problem of model repairment, while subsuming the existing works on influence functions and data deletion as specific instances. Experimentally, the proposed approach outperforms the baselines for both identification of detrimental training data and fixing model failures in a generalisable manner.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Ablating Concepts in Text-to-Image Diffusion ModelsNupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman et al.ICCV 2023 · 327 citations
- Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction AttacksVaidehi Patil, Peter Hase, Mohit BansalICLR 2024 · 167 citations
- Data Attribution for Text-to-Image Models by Unlearning Synthesized ImagesSheng-Yu Wang, Aaron Hertzmann, Alexei A. Efros, Jun-Yan Zhu et al.NeurIPS 2024 · 28 citations
- The Memory-Perturbation Equation: Understanding Model's Sensitivity to DataPeter Nickl, Lu Xu, Dharmesh Tailor, Thomas Möllenhoff et al.NeurIPS 2023 · 17 citations
- Redirection for Erasing Memory (REM): Towards a universal unlearning method for corrupted dataStefan Schoepf, Michael Mozer, Nicole Mitchell, Alexandra Brintrup et al.ICLR 2026 · 7 citations
Builds on10
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
- Adaptive Machine UnlearningVarun Gupta, Christopher Jung, Seth Neel, Aaron Roth et al.NeurIPS 2021 · 262 citations
- Editable Neural NetworksAnton Sinitsin, Vsevolod Plokhotnyuk, Dmitry V. Pyrkin, Sergei Popov et al.ICLR 2020 · 210 citations
- Variational Bayesian UnlearningQuoc Phong Nguyen, Bryan Kian Hsiang Low, Patrick JailletNeurIPS 2020 · 198 citations
Related papers
- Retaining Beneficial Information from Detrimental Data for Neural Network RepairLong-Kai Huang, Peilin Zhao, Junzhou Huang, Sinno Jialin PanNeurIPS 2023 · 2 citations
- Data Glitches Discovery using Influence-based Model ExplanationsNikolaos Myrtakis, Ioannis Tsamardinos, Vassilis ChristophidesKDD 2025
- ErrorEraser: Unlearning Data Bias for Improved Continual LearningXuemei Cao, Hanlin Gu, Xin Yang, Bingjun Wei et al.KDD 2025 · 1 citation
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
- Machine Unlearning of Features and LabelsAlexander Warnecke, Lukas Pirch, Christian Wressnegger, Konrad RieckNDSS 2023
