Enhancing Noise-Robust Losses for Large-Scale Noisy Data Learning
Max Staats, Matthias Thamm, Bernd Rosenow
Abstract
Large annotated datasets inevitably contain noisy labels, which poses a major challenge for training deep neural networks as they easily memorize the labels. Noise-robust loss functions have emerged as a notable strategy to counteract this issue, but it remains challenging to create a robust loss function which is not susceptible to underfitting. Through a quantitative approach, this paper explores the limited overlap between the network output at initialization and regions of non-vanishing gradients of bounded loss functions in the initial learning phase. Using these insights, we address underfitting of several noise robust losses with a novel method denoted as logit bias, which adds a real number epsilon to the logit at the position of the correct class. The logit bias enables these losses to achieve state-of-the-art results, even on datasets like WebVision, consisting of over a million images from 1000 classes. In addition, we demonstrate that our method can be used to determine optimal parameters for several loss functions – without having to train networks. Remarkably, our method determines the hyperparameters based on the number of classes, resulting in loss functions which require zero dataset or noise-dependent parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Normalized Loss Functions for Deep Learning with Noisy LabelsXingjun Ma, Hanxun Huang, Yisen Wang, Simone Romano et al.ICML 2020 · 547 citations
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor et al.NeurIPS 2021 · 208 citations
- Curriculum Loss: Robust Learning and Generalization against Label CorruptionYueming Lyu, Ivor W. TsangICLR 2020 · 190 citations
- Generalized Jensen-Shannon Divergence Loss for Learning with Noisy LabelsErik Englesson, Hossein AzizpourNeurIPS 2021 · 170 citations
Related papers
- Mitigating Memorization of Noisy Labels by Clipping the Model PredictionHongxin Wei, Huiping Zhuang, Renchunzi Xie, Lei Feng et al.ICML 2023 · 54 citations
- Correlated Input-Dependent Label Noise in Large-Scale Image ClassificationMark Collier, Basil Mustafa, Efi Kokiopoulou, Rodolphe Jenatton et al.CVPR 2021
- Sample-wise Label Confidence Incorporation for Learning with Noisy LabelsChanho Ahn, Kikyung Kim, Ji-Won Baek, Jongin Lim et al.ICCV 2023 · 11 citations
- Learning from Noisy Labels with Complementary Loss FunctionsDeng-Bao Wang, Yong Wen, Lujia Pan, Min-Ling ZhangAAAI 2021 · 40 citations
- O2U-Net: A Simple Noisy Label Detection Approach for Deep Neural NetworksJinchi Huang, Lie Qu, Rongfei Jia, Binqiang ZhaoICCV 2019 · 276 citations
