Sample Selection via Contrastive Fragmentation for Noisy Label Regression
Chris Dongjoo Kim, Sangwoo Moon, Jihwan Moon, Dongyeon Woo, Gunhee Kim
Abstract
As with many other problems, real-world regression is plagued by the presence of noisy labels, an inevitable issue that demands our attention. Fortunately, much real-world data often exhibits an intrinsic property of continuously ordered correlations between labels and features, where data points with similar labels are also represented with closely related features. In response, we propose a novel approach named ConFrag, where we collectively model the regression data by transforming them into disjoint yet contrasting fragmentation pairs. This enables the training of more distinctive representations, enhancing the ability to select clean samples. Our ConFrag framework leverages a mixture of neighboring fragments to discern noisy labels through neighborhood agreement among expert feature extractors. We extensively perform experiments on six newly curated benchmark datasets of diverse domains, including age prediction, price prediction, and music production year estimation. We also introduce a metric called Error Residual Ratio (ERR) to better account for varying degrees of label noise. Our approach consistently outperforms fourteen state-of-the-art baselines, being robust against symmetric and random Gaussian label noise.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78c30b83-667a-4b2f-b206-d8dc669e3722Cited by top-tier papers4
- Denoising Mixup for RegressionZhengzhang Hou, Zhanshan Li, Yanbo Liu, Geoff Nitschke et al.AAAI 2026
- Fault Lines: Benchmarking the Impact of Label Data Quality on ML Robustness and FairnessDavid Jackson, Paul Groth, Hazar HarmouchVLDB 2026
- Evolutionary Multi-View Classification with Label Noise via Gradient and Feature Dual-PerceptionShuai Li, Xinyan Liang, Yuhua Qian, Li LvICML 2026
- Stochastic Order Learning: An Approach to Rank Estimation Using Noisy DataChaewon Lee, Seon-Ho Lee, Chang-Su KimICML 2026
Builds on50
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- Connecting the Dots: Multivariate Time Series Forecasting with Graph Neural NetworksZonghan Wu, Shirui Pan, Guodong Long, Jing Jiang et al.KDD 2020 · 1,738 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo et al.ICCV 2019 · 1,125 citations
Related papers
- Enhancing Contrastive Learning with Noise-Guided Attack: Towards Continual Relation Extraction in the WildTing Wu, Jingyi Liu, Rui Zheng, Tao Gui et al.ACL 2024
- Contrastive Order Learning: A General Framework for Ordinal RegressionChaewon Lee, BeomJun Shim, Kwang Choi, Chang-Su KimICML 2026
- DAT: Training Deep Networks Robust To Label-Noise by Matching the Feature DistributionsYuntao Qu, Shasha Mo, Jianwei NiuCVPR 2021
- Semi-Supervised Contrastive Learning for Deep Regression with Ordinal Rankings from Spectral SeriationWeihang Dai, Yao Du, Hanru Bai, Kwang-Ting Cheng et al.NeurIPS 2023 · 16 citations
- Appearance Contrasts for Unconstrained Age EstimationJilong Wei, Yangyang Hu, Xiangjuan Wu, Yiqiang Wu et al.ACM MM 2025
