How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
Wei Huang, Andi Han, Yujin Song, Yilan Chen, Denny Wu, Difan Zou, Taiji Suzuki
Abstract
The capacity of deep learning models is often large enough to both learn the underlying statistical signal and overfit to noise in the training set. This noise memorization can be harmful especially for data with a low signal-to-noise ratio (SNR), leading to poor generalization. Inspired by prior observations that label noise provides implicit regularization that improves generalization, in this work, we investigate whether introducing label noise to the gradient updates can enhance the test performance of neural network (NN) in the low SNR regime. Specifically, we consider training a two-layer NN with a simple label noise gradient descent (GD) algorithm, in an idealized signal-noise data setting. We prove that adding label noise during training suppresses noise memorization, preventing it from dominating the learning process; consequently, label noise GD enjoys rapid signal growth while the overfitting remains controlled, thereby achieving good generalization despite the low SNR. In contrast, we also show that NN trained with standard GD tends to overfit to noise in the same low SNR setting and establish a non-vanishing lower bound on its test error, thus demonstrating the benefit of introducing label noise in gradient-based training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- When Do Neural Networks Outperform Kernel Methods?Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, Andrea MontanariNeurIPS 2020 · 217 citations
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang et al.NeurIPS 2022 · 173 citations
- Toward Understanding the Feature Learning Process of Self-supervised Contrastive LearningZixin Wen, Yuanzhi LiICML 2021 · 162 citations
Related papers
- Why Does Sharpness-Aware Minimization Generalize Better Than SGD?Zixiang Chen, Junkai Zhang, Yiwen Kou, Xiangning Chen et al.NeurIPS 2023 · 32 citations
- On the Role of Label Noise in the Feature Learning ProcessAndi Han, Wei Huang, Zhanpeng Zhou, Gang Niu et al.ICML 2025
- Towards Understanding Generalization in DP-GD: A Case Study in Training Two-Layer CNNsZhongjie Shi, Puyu Wang, Chenyang Zhang, Yuan CaoAAAI 2026 · 3 citations
- SIGUA: Forgetting May Make Learning with Noisy Labels More RobustBo Han, Gang Niu, Xingrui Yu, Quanming Yao et al.ICML 2020 · 157 citations
- Understanding Why Neural Networks Generalize Well Through GSNR of ParametersJinlong Liu, Yunzhi Bai, Guoqing Jiang, Ting Chen et al.ICLR 2020 · 60 citations
