Noise Attention Learning: Enhancing Noise Robustness by Gradient Scaling
Yangdi Lu, Yang Bo, Wenbo He
Abstract
Machine learning has been highly successful in data-driven applications but is often hampered when the data contains noise, especially label noise. When trained on noisy labels, deep neural networks tend to fit all noisy labels, resulting in poor generalization. To handle this problem, a common idea is to force the model to fit only clean samples rather than mislabeled ones. In this paper, we propose a simple yet effective method that automatically distinguishes the mislabeled samples and prevents the model from memorizing them, named Noise Attention Learning. In our method, we introduce an attention branch to produce attention weights based on representations of samples. This attention branch is learned to divide the samples according to the predictive power in their representations. We design the corresponding loss function that incorporates the attention weights for training the model without affecting the original learning direction. Empirical results show that most of the mislabeled samples yield significantly lower weights than the clean ones. Furthermore, our theoretical analysis shows that the gradients of training samples are dynamically scaled by the attention weights, implicitly preventing memorization of the mislabeled samples. Experimental results on two benchmarks (CIFAR-10 and CIFAR-100) with simulated label noise and three realworld noisy datasets (ANIMAL-10N, Clothing1M and Webvision) demonstrate that our approach outperforms state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Robust Classification via Regression for Learning with Noisy LabelsErik Englesson, Hossein AzizpourICLR 2024 · 12 citations
- Theoretically Guaranteed Bidirectional Data Rectification for Robust Sequential RecommendationYatong Sun, Bin Wang, Zhu Sun, Xiaochun Yang et al.NeurIPS 2023 · 11 citations
- Subclass-Dominant Label Noise: A Counterexample for the Success of Early StoppingYingbin Bai, Zhongyi Han, Erkun Yang, Jun Yu et al.NeurIPS 2023 · 10 citations
- Revisiting Interpolation for Noisy Label CorrectionYuanzhuo Xu, Xiaoguang Niu, Jie Yang, Ruiyi Su et al.AAAI 2025 · 8 citations
- Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic OptimizationKuan Zhang, Chengliang Chai, Jingzhe Xu, Chi Zhang et al.NeurIPS 2025 · 6 citations
Builds on20
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo et al.ICCV 2019 · 1,125 citations
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 798 citations
- Normalized Loss Functions for Deep Learning with Noisy LabelsXingjun Ma, Hanxun Huang, Yisen Wang, Simone Romano et al.ICML 2020 · 547 citations
Related papers
- On the Role of Label Noise in the Feature Learning ProcessAndi Han, Wei Huang, Zhanpeng Zhou, Gang Niu et al.ICML 2025
- Jo-SRC: A Contrastive Approach for Combating Noisy LabelsYazhou Yao, Zeren Sun, Chuanyi Zhang, Fumin Shen et al.CVPR 2021
- DAT: Training Deep Networks Robust To Label-Noise by Matching the Feature DistributionsYuntao Qu, Shasha Mo, Jianwei NiuCVPR 2021
- Enhancing Robustness in Learning with Noisy Labels: An Asymmetric Co-Training ApproachMengmeng Sheng, Zeren Sun, Gensheng Pei, Tao Chen et al.ACM MM 2024 · 7 citations
- Large Loss Matters in Weakly Supervised Multi-Label ClassificationYoungwook Kim, Jae-Myung Kim, Zeynep Akata, Jungwoo LeeCVPR 2022 · 68 citations
