Retaining Beneficial Information from Detrimental Data for Neural Network Repair
Long-Kai Huang, Peilin Zhao, Junzhou Huang, Sinno Jialin Pan
摘要
The performance of deep learning models heavily relies on the quality of the training data. Inadequacies in the training data, such as corrupt input or noisy labels, can lead to the failure of model generalization. Recent studies propose repairing the model by identifying the training samples that contribute to the failure and removing their influence from the model. However, it is important to note that the identified data may contain both beneficial and detrimental information. Simply erasing the information of the identified data from the model can have a negative impact on its performance, especially when accurate data is mistakenly identified as detrimental and removed. To overcome this challenge, we propose a novel approach that leverages the knowledge obtained from a retained clean set. Concretely, Our method first identifies harmful data by utilizing the clean set, then separates the beneficial and detrimental information within the identified data. Finally, we utilize the extracted beneficial information to enhance the model’s performance. Through empirical evaluations, we demonstrate that our method outperforms baseline approaches in both identifying harmful data and rectifying model failures. Particularly in scenarios where identification is challenging and a significant amount of benign data is involved, our method improves performance while the baselines deteriorate due to the erroneous removal of beneficial information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
- Model Editing for Vision TransformersXinyi Huang, Kangfei Zhao, Long-Kai HuangNeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper12
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud 等ICLR 2020 · 被引用 643 次
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 被引用 633 次
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu 等ICLR 2022 · 被引用 338 次
- The Early Phase of Neural Network TrainingJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2020 · 被引用 199 次
相关 Paper
- Repairing Neural Networks by Leaving the Right Past BehindRyutaro Tanno, Melanie F. Pradier, Aditya V. Nori, Yingzhen LiNeurIPS 2022 · 被引用 41 次
- Learning Deep Neural Networks under Agnostic Corrupted SupervisionBoyang Liu, Mengying Sun, Ding Wang, Pang-Ning Tan 等ICML 2021 · 被引用 7 次
- A Simple Remedy for Dataset Bias via Self-Influence: A Mislabeled Sample PerspectiveYeonsung Jung, Jaeyun Song, June Yong Yang, Jin-Hwa Kim 等NeurIPS 2024 · 被引用 7 次
- Differential Training: A Generic Framework to Reduce Label Noises for Android Malware DetectionJiayun Xu, Yingjiu Li, Robert H. DengNDSS 2021
- Dynamic Data Fault Localization for Deep Neural NetworksYining Yin, Yang Feng, Shihao Weng, Zixi Liu 等FSE 2023 · 被引用 10 次
