Robust Curriculum Learning: from clean label detection to noisy label self-correction
Tianyi Zhou, Shengjie Wang, Jeff A. Bilmes
Abstract
Neural network training can easily overfit noisy labels resulting in poor generalization performance. Existing methods address this problem by (1) filtering out the noisy data and only using the clean data for training or (2) relabeling the noisy data by the model during training or by another model trained only on a clean dataset. However, the former does not leverage the features' information of wrongly-labeled data, while the latter may produce wrong pseudo-labels for some data and introduce extra noises. In this paper, we propose a smooth transition and interplay between these two strategies as a curriculum that selects training samples dynamically. In particular, we start with learning from clean data and then gradually move to learn noisy-labeled data with pseudo labels produced by a time-ensemble of the model and data augmentations. Instead of using the instantaneous loss computed at the current step, our data selection is based on the dynamics of both the loss and output consistency for each sample across historical steps and different data augmentations, resulting in more precise detection of both clean labels and correct pseudo labels. On multiple benchmarks of noisy labels, we show that our curriculum learning strategy can significantly improve the test accuracy without any auxiliary model or extra clean data.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers36
- Adaptive Early-Learning Correction for Segmentation from Noisy AnnotationsSheng Liu, Kangning Liu, Weicheng Zhu, Yiqiu Shen et al.CVPR 2022 · 109 citations
- Combating Noisy Labels with Sample Selection by Mining High-Discrepancy ExamplesXiaobo Xia, Bo Han, Yibing Zhan, Jun Yu et al.ICCV 2023 · 72 citations
- Sketching without Worrying: Noise-Tolerant Sketch-Based Image RetrievalAyan Kumar Bhunia, Subhadeep Koley, Abdullah Faiz Ur Rahman Khilji, Aneeshan Sain et al.CVPR 2022 · 53 citations
- Robust Data Pruning under Label Noise via Maximizing Re-labeling AccuracyDongmin Park, Seola Choi, Doyoung Kim, Hwanjun Song et al.NeurIPS 2023 · 42 citations
- Early Stopping Against Label Noise Without Validation DataSuqin Yuan, Lei Feng, Tongliang LiuICLR 2024 · 39 citations
Related papers
- SuperLoss: A Generic Loss for Robust Curriculum LearningThibault Castells, Philippe Weinzaepfel, Jérôme RevaudNeurIPS 2020 · 96 citations
- SELF: Learning to Filter Noisy Labels with Self-EnsemblingDuc Tam Nguyen, Chaithanya Kumar Mummadi, Thi-Phuong-Nhung Ngo, Thi Hoai Phuong Nguyen et al.ICLR 2020 · 354 citations
- Time-Consistent Self-Supervision for Semi-Supervised LearningTianyi Zhou, Shengjie Wang, Jeff A. BilmesICML 2020 · 58 citations
- TrainRef: Curating Data with Label Distribution and Minimal Reference for Accurate Prediction and Reliable ConfidenceMurong Ma, Ruofan Liu, Yun Lin, Zhiyong Huang et al.ICLR 2026
- C-SFDA: A Curriculum Learning Aided Self-Training Framework for Efficient Source Free Domain AdaptationNazmul Karim, Niluthpol Chowdhury Mithun, Abhinav Rajvanshi, Han-Pang Chiu et al.CVPR 2023
