Retaining Beneficial Information from Detrimental Data for Neural Network Repair
Long-Kai Huang, Peilin Zhao, Junzhou Huang, Sinno Jialin Pan
Abstract
The performance of deep learning models heavily relies on the quality of the training data. Inadequacies in the training data, such as corrupt input or noisy labels, can lead to the failure of model generalization. Recent studies propose repairing the model by identifying the training samples that contribute to the failure and removing their influence from the model. However, it is important to note that the identified data may contain both beneficial and detrimental information. Simply erasing the information of the identified data from the model can have a negative impact on its performance, especially when accurate data is mistakenly identified as detrimental and removed. To overcome this challenge, we propose a novel approach that leverages the knowledge obtained from a retained clean set. Concretely, Our method first identifies harmful data by utilizing the clean set, then separates the beneficial and detrimental information within the identified data. Finally, we utilize the extracted beneficial information to enhance the model’s performance. Through empirical evaluations, we demonstrate that our method outperforms baseline approaches in both identifying harmful data and rectifying model failures. Particularly in scenarios where identification is challenging and a significant amount of benign data is involved, our method improves performance while the baselines deteriorate due to the erroneous removal of beneficial information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f00ce533-7182-43cc-a42b-93c401541cb5Cited by top-tier papers1
Ask how each one uses itBuilds on12
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu et al.ICLR 2022 · 338 citations
- The Early Phase of Neural Network TrainingJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2020 · 199 citations
Related papers
- Repairing Neural Networks by Leaving the Right Past BehindRyutaro Tanno, Melanie F. Pradier, Aditya V. Nori, Yingzhen LiNeurIPS 2022 · 41 citations
- Learning Deep Neural Networks under Agnostic Corrupted SupervisionBoyang Liu, Mengying Sun, Ding Wang, Pang-Ning Tan et al.ICML 2021 · 7 citations
- A Simple Remedy for Dataset Bias via Self-Influence: A Mislabeled Sample PerspectiveYeonsung Jung, Jaeyun Song, June Yong Yang, Jin-Hwa Kim et al.NeurIPS 2024 · 7 citations
- Differential Training: A Generic Framework to Reduce Label Noises for Android Malware DetectionJiayun Xu, Yingjiu Li, Robert H. DengNDSS 2021
- Dynamic Data Fault Localization for Deep Neural NetworksYining Yin, Yang Feng, Shihao Weng, Zixi Liu et al.FSE 2023 · 10 citations
