CrossSplit: Mitigating Label Noise Memorization through Data Splitting
Jihye Kim, Aristide Baratin, Yan Zhang, Simon Lacoste-Julien
Abstract
We approach the problem of improving robustness of deep learning algorithms in the presence of label noise. Building upon existing label correction and co-teaching methods, we propose a novel training procedure to mitigate the memorization of noisy labels, called CrossSplit, which uses a pair of neural networks trained on two disjoint parts of the labelled dataset. CrossSplit combines two main ingredients: (i) Cross-split label correction. The idea is that, since the model trained on one part of the data cannot memorize example-label pairs from the other part, the training labels presented to each network can be smoothly adjusted by using the predictions of its peer network; (ii) Cross-split semi-supervised training. A network trained on one part of the data also uses the unlabeled inputs of the other part. Extensive experiments on CIFAR-10, CIFAR-100, Tiny-ImageNet and mini-WebVision datasets demonstrate that our method can outperform the current state-of-the-art in a wide range of noise ratios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50305cd6-1573-47e0-b049-b8be2c36dfaeCited by top-tier papers5
- Discovering Environments with XRMMohammad Pezeshki, Diane Bouchacourt, Mark Ibrahim, Nicolas Ballas et al.ICML 2024 · 21 citations
- CLIPCleaner: Cleaning Noisy Labels with CLIPChen Feng, Georgios Tzimiropoulos, Ioannis PatrasACM MM 2024 · 12 citations
- Revisiting Interpolation for Noisy Label CorrectionYuanzhuo Xu, Xiaoguang Niu, Jie Yang, Ruiyi Su et al.AAAI 2025 · 8 citations
- Combating Semantic Contamination in Learning with Label NoiseWenxiao Fan, Kan LiAAAI 2025 · 1 citation
- Leveraging Dissimilarity Invariance as a Robust Anchor for Learning with Noisy LabelsWenxiao Fan, Kan LiAAAI 2026
Builds on8
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 798 citations
- Normalized Loss Functions for Deep Learning with Noisy LabelsXingjun Ma, Hanxun Huang, Yisen Wang, Simone Romano et al.ICML 2020 · 547 citations
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 398 citations
- Deep Learning Through the Lens of Example DifficultyRobert J. N. Baldock, Hartmut Maennel, Behnam NeyshaburNeurIPS 2021 · 204 citations
- Selective-Supervised Contrastive Learning with Noisy LabelsShikun Li, Xiaobo Xia, Shiming Ge, Tongliang LiuCVPR 2022 · 201 citations
Related papers
- Enhancing Robustness in Learning with Noisy Labels: An Asymmetric Co-Training ApproachMengmeng Sheng, Zeren Sun, Gensheng Pei, Tao Chen et al.ACM MM 2024 · 7 citations
- Combating Noisy Labels by Agreement: A Joint Training Method with Co-RegularizationHongxin Wei, Lei Feng, Xiangyu Chen, Bo AnCVPR 2020
- Learning from Noisy Labels with Complementary Loss FunctionsDeng-Bao Wang, Yong Wen, Lujia Pan, Min-Ling ZhangAAAI 2021 · 40 citations
- Mitigating Label Noise through Data AmbiguationJulian Lienen, Eyke HüllermeierAAAI 2024 · 14 citations
- Coresets for Robust Training of Deep Neural Networks against Noisy LabelsBaharan Mirzasoleiman, Kaidi Cao, Jure LeskovecNeurIPS 2020 · 99 citations
