Robust Classification via Regression for Learning with Noisy Labels
Erik Englesson, Hossein Azizpour
Abstract
Deep neural networks and large-scale datasets have revolutionized the field of machine learning. However, these large networks are susceptible to overfitting to label noise, resulting in reduced generalization. To address this challenge, two promising approaches have emerged: i) loss reweighting, which reduces the influence of noisy examples on the training loss, and ii) label correction that replaces noisy labels with estimated true labels. These directions have been pursued separately or combined as independent methods, lacking a unified approach. In this work, we present a unified method that seamlessly combines loss reweighting and label correction to enhance robustness against label noise in classification tasks. Specifically, by leveraging ideas from compositional data analysis in statistics, we frame the problem as a regression task, where loss reweighting and label correction can naturally be achieved with a shifted Gaussian label noise model. Our unified approach achieves strong performance compared to recent baselines on several noisy labelled datasets. We believe this work is a promising step towards robust deep learning in the presence of label noise. Our code is available at: github.com/ErikEnglesson/SGN.
Published as a conference paper at ICLR 2024 These directions have been pursued separately or combined as independent methods, lacking a unified approach. In this work, we present a unified method that seamlessly combines loss reweighting and label correction to enhance robustness against label noise in classification tasks. More precisely, our main contributions are:
• We propose an adaptation of the log-ratio transform approach from compositional data analysis in statistics to the classification task (Section 3.1). This includes turning the classification dataset to a regression dataset (Section 3.2) and to transform regression predictions to classification predictions (Section 3.4).
• With this novel view, we solve the label noise problem in classification as a regression problem. We present a unified probabilistic regression method that seamlessly combines loss reweighting and label correction (Section 3.3). We naturally achieve loss reweighting by learning the mean and covariance of per-example Gaussian distributions. To achieve label correction, we naturally extend this approach by using a shifted (non-zero mean) Gaussian noise model, which we show indirectly changes the label.
• Finally, we perform extensive experiments, that increases our understanding of the method and shows its strong performance compared to baselines on several datasets (Section 4).
In this section, we provide background information about the statistical field of compositional data analysis (Aitchison, 2011), and in particular the log-ratio transform approach.
Compositional Data Compositional data is a collection of compositions, which are non-negative real vectors that sum to a constant, usually 1. Indeed, in this case, compositions are categorical distributions in the probability simplex
where D denotes the number of components. Compositional data naturally arises in many fields of study (Aitchison, 2005;Tsagris & Stewart, 2020), e.g., in geology when studying [sand, silt, clay] compositions of sedimentary rock. As compositions are constrained variables, compositional data cannot be analysed with statistical techniques designed for unconstrained variables (Pawlowsky-Glahn & Egozcue, 2006). Next, we describe a suitable method for analysing compositional data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3c31b99-7fa1-4e7c-81db-0b995c69e97fCited by top-tier papers3
- Enhancing Sample Selection Against Label Noise by Cutting Mislabeled Easy ExamplesSuqin Yuan, Lei Feng, Bo Han, Tongliang LiuNeurIPS 2025 · 5 citations
- On Revisiting Entropy for Identifying Mislabeled ImagesChunlei Li, Zixuan Zheng, Yilei Shi, Guanglu Dong et al.ICML 2026
- Just Y-Prediction: Enabling Historical Cumulative Inconsistency in Label Diffusion for Learning with Noisy LabelSenyu Hou, Gaoxia Jiang, Xinyi Zheng, Yaqing Guo et al.ICML 2026
Builds on18
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo et al.ICCV 2019 · 1,125 citations
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 798 citations
- Normalized Loss Functions for Deep Learning with Noisy LabelsXingjun Ma, Hanxun Huang, Yisen Wang, Simone Romano et al.ICML 2020 · 547 citations
Related papers
- Training Noise-Robust Deep Neural Networks via Meta-LearningZhen Wang, Guosheng Hu, Qinghua HuCVPR 2020
- Mitigating Memorization of Noisy Labels by Clipping the Model PredictionHongxin Wei, Huiping Zhuang, Renchunzi Xie, Lei Feng et al.ICML 2023 · 54 citations
- Tackling Instance-Dependent Label Noise with Class Rebalance and Geometric RegularizationShuzhi Cao, Jianfei Ruan, Bo Dong, Bin ShiKDD 2024 · 1 citation
- Distilling Effective Supervision From Severe Label NoiseZizhao Zhang, Han Zhang, Sercan Ömer Arik, Honglak Lee et al.CVPR 2020
- L2B: Learning to Bootstrap Robust Models for Combating Label NoiseYuyin Zhou, Xianhang Li, Fengze Liu, Qingyue Wei et al.CVPR 2024 · 13 citations
