Combating Noisy Labels with Sample Selection by Mining High-Discrepancy Examples
Xiaobo Xia, Bo Han, Yibing Zhan, Jun Yu, Mingming Gong, Chen Gong, Tongliang Liu
Abstract
The sample selection approach is popular in learning with noisy labels. The state-of-the-art methods train two deep networks simultaneously for sample selection, which aims to employ their different learning abilities. To prevent two networks from converging to a consensus, their divergence should be maintained. Prior work presents that the divergence can be kept by locating the disagreement data on which the prediction labels of the two networks are different. However, this procedure is sample-inefficient for generalization, which means that only a few clean examples can be utilized in training. In this paper, to address the issue, we propose a simple yet effective method called CoDis. In particular, we select possibly clean data that simultaneously have high-discrepancy prediction probabilities between two networks. As selected data have high discrepancies in probabilities, the divergence of two networks can be maintained by training on such data. In addition, the condition of high discrepancies is milder than disagreement, which allows more data to be considered for training, and makes our method more sample-efficient. Moreover, we show that the proposed method enables to mine hard clean examples to help generalization. Empirical results show that CoDis is superior to multiple baselines in the robustness of trained models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- FlatMatch: Bridging Labeled Data and Unlabeled Data with Cross-Sharpness for Semi-Supervised LearningZhuo Huang, Li Shen, Jun Yu, Bo Han et al.NeurIPS 2023 · 50 citations
- CLIPLoss and Norm-Based Data Selection Methods for Multimodal Contrastive LearningYiping Wang, Yifang Chen, Wendan Yan, Alex Fang et al.NeurIPS 2024 · 31 citations
- On the Over-Memorization During Natural, Robust and Catastrophic OverfittingRunqi Lin, Chaojian Yu, Bo Han, Tongliang LiuICLR 2024 · 21 citations
- Unlocking the Power of Open Set: A New Perspective for Open-Set Noisy Label LearningWenhai Wan, Xinrui Wang, Ming-Kun Xie, Shao-Yuan Li et al.AAAI 2024 · 18 citations
- Curriculum Fine-tuning of Vision Foundation Model for Medical Image Classification Under Label NoiseYeonguk Yu, Minhwan Ko, Sungho Shin, Kangmin Kim et al.NeurIPS 2024 · 10 citations
Builds on35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 798 citations
- Normalized Loss Functions for Deep Learning with Noisy LabelsXingjun Ma, Hanxun Huang, Yisen Wang, Simone Romano et al.ICML 2020 · 547 citations
Related papers
- Enhancing Robustness in Learning with Noisy Labels: An Asymmetric Co-Training ApproachMengmeng Sheng, Zeren Sun, Gensheng Pei, Tao Chen et al.ACM MM 2024 · 7 citations
- Jo-SRC: A Contrastive Approach for Combating Noisy LabelsYazhou Yao, Zeren Sun, Chuanyi Zhang, Fumin Shen et al.CVPR 2021
- Combating Noisy Labels by Agreement: A Joint Training Method with Co-RegularizationHongxin Wei, Lei Feng, Xiangyu Chen, Bo AnCVPR 2020
- PNP: Robust Learning from Noisy Labels by Probabilistic Noise PredictionZeren Sun, Fumin Shen, Dan Huang, Qiong Wang et al.CVPR 2022 · 79 citations
- DAT: Training Deep Networks Robust To Label-Noise by Matching the Feature DistributionsYuntao Qu, Shasha Mo, Jianwei NiuCVPR 2021
