Lune

ICLR2024顶会

Robust Classification via Regression for Learning with Noisy Labels

Erik Englesson, Hossein Azizpour

出版方
2024年份
12被引次数
3顶会引用

摘要

Deep neural networks and large-scale datasets have revolutionized the field of machine learning. However, these large networks are susceptible to overfitting to label noise, resulting in reduced generalization. To address this challenge, two promising approaches have emerged: i) loss reweighting, which reduces the influence of noisy examples on the training loss, and ii) label correction that replaces noisy labels with estimated true labels. These directions have been pursued separately or combined as independent methods, lacking a unified approach. In this work, we present a unified method that seamlessly combines loss reweighting and label correction to enhance robustness against label noise in classification tasks. Specifically, by leveraging ideas from compositional data analysis in statistics, we frame the problem as a regression task, where loss reweighting and label correction can naturally be achieved with a shifted Gaussian label noise model. Our unified approach achieves strong performance compared to recent baselines on several noisy labelled datasets. We believe this work is a promising step towards robust deep learning in the presence of label noise. Our code is available at: github.com/ErikEnglesson/SGN.

Published as a conference paper at ICLR 2024 These directions have been pursued separately or combined as independent methods, lacking a unified approach. In this work, we present a unified method that seamlessly combines loss reweighting and label correction to enhance robustness against label noise in classification tasks. More precisely, our main contributions are:

• We propose an adaptation of the log-ratio transform approach from compositional data analysis in statistics to the classification task (Section 3.1). This includes turning the classification dataset to a regression dataset (Section 3.2) and to transform regression predictions to classification predictions (Section 3.4).

• With this novel view, we solve the label noise problem in classification as a regression problem. We present a unified probabilistic regression method that seamlessly combines loss reweighting and label correction (Section 3.3). We naturally achieve loss reweighting by learning the mean and covariance of per-example Gaussian distributions. To achieve label correction, we naturally extend this approach by using a shifted (non-zero mean) Gaussian noise model, which we show indirectly changes the label.

• Finally, we perform extensive experiments, that increases our understanding of the method and shows its strong performance compared to baselines on several datasets (Section 4).

In this section, we provide background information about the statistical field of compositional data analysis (Aitchison, 2011), and in particular the log-ratio transform approach.

Compositional Data Compositional data is a collection of compositions, which are non-negative real vectors that sum to a constant, usually 1. Indeed, in this case, compositions are categorical distributions in the probability simplex

where D denotes the number of components. Compositional data naturally arises in many fields of study (Aitchison, 2005;Tsagris & Stewart, 2020), e.g., in geology when studying [sand, silt, clay] compositions of sedimentary rock. As compositions are constrained variables, compositional data cannot be analysed with statistical techniques designed for unconstrained variables (Pawlowsky-Glahn & Egozcue, 2006). Next, we describe a suitable method for analysing compositional data.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper18

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖