Delving into Noisy Label Detection with Clean Data
Chenglin Yu, Xinsong Ma, Weiwei Liu
Abstract
A critical element of learning with noisy labels is noisy label detection. Notably, numerous previous works assume that no source of labels can be clean in a noisy label detection context. In this work, we relax this assumption and assume that a small subset of the training data is clean, which enables substantial noisy label detection performance gains. Specifically, we propose a novel framework that leverages clean data by framing the problem of noisy label detection with clean data as a multiple hypothesis testing problem. Moreover, we propose BHN, a simple yet effective approach for noisy label detection that integrates the Benjamini-Hochberg (BH) procedure into deep neural networks. BHN achieves state-ofthe-art performance and outperforms baselines by 28.48% in terms of false discovery rate (FDR) and by 18.99% in terms of F1 on CIFAR-10. Extensive ablation studies further demonstrate the superiority of BHN. Our code is available at https://github.com/ChenglinYu/BHN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- VDC: Versatile Data Cleanser based on Visual-Linguistic Inconsistency by Multimodal Large Language ModelsZihao Zhu, Mingda Zhang, Shaokui Wei, Bingzhe Wu et al.ICLR 2024 · 11 citations
- A Provable Decision Rule for Out-of-Distribution DetectionXinsong Ma, Xin Zou, Weiwei LiuICML 2024 · 4 citations
- The Reliability of OKRidge Method in Solving Sparse Ridge Regression ProblemsXiyuan Li, Youjun Wang, Weiwei LiuNeurIPS 2024
- MISF: MLLM Guided Iterative Sample Filtering for Data Fault DetectionGuoying Chen, Ruizhuo Zhao, Zhewei Xu, Bo Yang et al.AAAI 2026
- A Closer Look at Generalized BH Algorithm for Out-of-Distribution DetectionXinsong Ma, Jie Wu, Weiwei LiuICML 2025
Builds on18
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo et al.ICCV 2019 · 1,125 citations
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 798 citations
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
- Normalized Loss Functions for Deep Learning with Noisy LabelsXingjun Ma, Hanxun Huang, Yisen Wang, Simone Romano et al.ICML 2020 · 547 citations
- Mitigating Neural Network Overconfidence with Logit NormalizationHongxin Wei, Renchunzi Xie, Hao Cheng, Lei Feng et al.ICML 2022 · 386 citations
Related papers
- Error-Bounded Correction of Noisy LabelsSongzhu Zheng, Pengxiang Wu, Aman Goswami, Mayank Goswami et al.ICML 2020 · 153 citations
- Scalable Penalized Regression for Noise Detection in Learning with Noisy LabelsYikai Wang, Xinwei Sun, Yanwei FuCVPR 2022 · 37 citations
- Sample-wise Label Confidence Incorporation for Learning with Noisy LabelsChanho Ahn, Kikyung Kim, Ji-Won Baek, Jongin Lim et al.ICCV 2023 · 11 citations
- FINE Samples for Learning with Noisy LabelsTaehyeon Kim, Jongwoo Ko, Sangwook Cho, Jinhwan Choi et al.NeurIPS 2021 · 145 citations
- DAT: Training Deep Networks Robust To Label-Noise by Matching the Feature DistributionsYuntao Qu, Shasha Mo, Jianwei NiuCVPR 2021
