Coresets for Robust Training of Deep Neural Networks against Noisy Labels
Baharan Mirzasoleiman, Kaidi Cao, Jure Leskovec
摘要
Modern neural networks have the capacity to overfit noisy labels frequently found in real-world datasets. Although great progress has been made, existing techniques are limited in providing theoretical guarantees for the performance of the neural networks trained with noisy labels. Here we propose a novel approach with strong theoretical guarantees for robust training of deep networks trained with noisy labels. The key idea behind our method is to select weighted subsets (coresets) of clean data points that provide an approximately low-rank Jacobian matrix. We then prove that gradient descent applied to the subsets do not overfit the noisy labels. Our extensive experiments corroborate our theory and demonstrate that deep networks trained on our subsets achieve a significantly superior performance compared to state-of-the art, e.g., 6% increase in accuracy on CIFAR-10 with 80% noisy labels, and 7% increase in accuracy on mini Webvision 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- RETRIEVE: Coreset Selection for Efficient and Robust Semi-Supervised LearningKrishnaTeja Killamsetty, Xujiang Zhao, Feng Chen, Rishabh K. IyerNeurIPS 2021 · 被引用 115 次
- Learning Noise Transition Matrix from Only Noisy Labels via Total Variation RegularizationYivan Zhang, Gang Niu, Masashi SugiyamaICML 2021 · 被引用 107 次
- GCR: Gradient Coreset based Replay Buffer Selection for Continual LearningRishabh Tiwari, KrishnaTeja Killamsetty, Rishabh K. Iyer, Pradeep ShenoyCVPR 2022 · 被引用 102 次
- Data Pruning via Moving-one-Sample-outHaoru Tan, Sitong Wu, Fei Du, Yukang Chen 等NeurIPS 2023 · 被引用 91 次
- Investigating Why Contrastive Learning Benefits Robustness against Label NoiseYihao Xue, Kyle Whitecross, Baharan MirzasoleimanICML 2022 · 被引用 70 次
它引用的顶会 Paper3
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- Heteroskedastic and Imbalanced Deep Learning with Adaptive RegularizationKaidi Cao, Yining Chen, Junwei Lu, Nikos Aréchiga 等ICLR 2021 · 被引用 20 次
- Distilling Effective Supervision From Severe Label NoiseZizhao Zhang, Han Zhang, Sercan Ömer Arik, Honglak Lee 等CVPR 2020
相关 Paper
- Robust Training under Label Noise by Over-parameterizationSheng Liu, Zhihui Zhu, Qing Qu, Chong YouICML 2022 · 被引用 152 次
- Sample-wise Label Confidence Incorporation for Learning with Noisy LabelsChanho Ahn, Kikyung Kim, Ji-Won Baek, Jongin Lim 等ICCV 2023 · 被引用 11 次
- Data-Efficient Augmentation for Training Neural NetworksTian Yu Liu, Baharan MirzasoleimanNeurIPS 2022 · 被引用 12 次
- Deep Self-Learning From Noisy LabelsJiangfan Han, Ping Luo, Xiaogang WangICCV 2019 · 被引用 315 次
- Error-Bounded Correction of Noisy LabelsSongzhu Zheng, Pengxiang Wu, Aman Goswami, Mayank Goswami 等ICML 2020 · 被引用 153 次
