Coresets for Robust Training of Deep Neural Networks against Noisy Labels
Baharan Mirzasoleiman, Kaidi Cao, Jure Leskovec
Abstract
Modern neural networks have the capacity to overfit noisy labels frequently found in real-world datasets. Although great progress has been made, existing techniques are limited in providing theoretical guarantees for the performance of the neural networks trained with noisy labels. Here we propose a novel approach with strong theoretical guarantees for robust training of deep networks trained with noisy labels. The key idea behind our method is to select weighted subsets (coresets) of clean data points that provide an approximately low-rank Jacobian matrix. We then prove that gradient descent applied to the subsets do not overfit the noisy labels. Our extensive experiments corroborate our theory and demonstrate that deep networks trained on our subsets achieve a significantly superior performance compared to state-of-the art, e.g., 6% increase in accuracy on CIFAR-10 with 80% noisy labels, and 7% increase in accuracy on mini Webvision 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8919f2df-e0f4-48f5-bcd6-0c80ef4dc0f9Cited by top-tier papers33
- RETRIEVE: Coreset Selection for Efficient and Robust Semi-Supervised LearningKrishnaTeja Killamsetty, Xujiang Zhao, Feng Chen, Rishabh K. IyerNeurIPS 2021 · 115 citations
- Learning Noise Transition Matrix from Only Noisy Labels via Total Variation RegularizationYivan Zhang, Gang Niu, Masashi SugiyamaICML 2021 · 107 citations
- GCR: Gradient Coreset based Replay Buffer Selection for Continual LearningRishabh Tiwari, KrishnaTeja Killamsetty, Rishabh K. Iyer, Pradeep ShenoyCVPR 2022 · 102 citations
- Data Pruning via Moving-one-Sample-outHaoru Tan, Sitong Wu, Fei Du, Yukang Chen et al.NeurIPS 2023 · 91 citations
- Investigating Why Contrastive Learning Benefits Robustness against Label NoiseYihao Xue, Kyle Whitecross, Baharan MirzasoleimanICML 2022 · 70 citations
Builds on3
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Heteroskedastic and Imbalanced Deep Learning with Adaptive RegularizationKaidi Cao, Yining Chen, Junwei Lu, Nikos Aréchiga et al.ICLR 2021 · 20 citations
- Distilling Effective Supervision From Severe Label NoiseZizhao Zhang, Han Zhang, Sercan Ömer Arik, Honglak Lee et al.CVPR 2020
Related papers
- Robust Training under Label Noise by Over-parameterizationSheng Liu, Zhihui Zhu, Qing Qu, Chong YouICML 2022 · 152 citations
- Sample-wise Label Confidence Incorporation for Learning with Noisy LabelsChanho Ahn, Kikyung Kim, Ji-Won Baek, Jongin Lim et al.ICCV 2023 · 11 citations
- Data-Efficient Augmentation for Training Neural NetworksTian Yu Liu, Baharan MirzasoleimanNeurIPS 2022 · 12 citations
- Deep Self-Learning From Noisy LabelsJiangfan Han, Ping Luo, Xiaogang WangICCV 2019 · 315 citations
- Error-Bounded Correction of Noisy LabelsSongzhu Zheng, Pengxiang Wu, Aman Goswami, Mayank Goswami et al.ICML 2020 · 153 citations
