Detecting Corrupted Labels Without Training a Model to Predict
Zhaowei Zhu, Zihao Dong, Yang Liu
Abstract
Label noise in real-world datasets encodes wrong correlation patterns and impairs the generalization of deep neural networks (DNNs). It is critical to find efficient ways to detect corrupted patterns. Current methods primarily focus on designing robust training techniques to prevent DNNs from memorizing corrupted patterns. These approaches often require customized training processes and may overfit corrupted patterns, leading to a performance drop in detection. In this paper, from a more data-centric perspective, we propose a training-free solution to detect corrupted labels. Intuitively, closer'' instances are more likely to share the same clean label. Based on the neighborhood information, we propose two methods: the first one uses local voting"via checking the noisy label consensuses of nearby features. The second one is a ranking-based approach that scores each instance and filters out a guaranteed number of instances that are likely to be corrupted. We theoretically analyze how the quality of features affects the local voting and provide guidelines for tuning neighborhood size. We also prove the worst-case error bound for the ranking-based method. Experiments with both synthetic and real-world label noise demonstrate our training-free solutions consistently and significantly improve most of the training-based baselines. Code is available at github.com/UCSC-REAL/SimiFeat.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 522d084e-47dd-4f31-a8ea-787648abff60Cited by top-tier papers29
- Mitigating Neural Network Overconfidence with Logit NormalizationHongxin Wei, Renchunzi Xie, Hao Cheng, Lei Feng et al.ICML 2022 · 386 citations
- Mitigating Memorization of Noisy Labels by Clipping the Model PredictionHongxin Wei, Huiping Zhuang, Renchunzi Xie, Lei Feng et al.ICML 2023 · 54 citations
- Open-Sampling: Exploring Out-of-Distribution data for Re-balancing Long-tailed datasetsHongxin Wei, Lue Tao, Renchunzi Xie, Lei Feng et al.ICML 2022 · 46 citations
- The Rich Get Richer: Disparate Impact of Semi-Supervised LearningZhaowei Zhu, Tianyi Luo, Yang LiuICLR 2022 · 44 citations
- Beyond Images: Label Noise Transition Matrix Estimation for Tasks with Lower-Quality FeaturesZhaowei Zhu, Jialu Wang, Yang LiuICML 2022 · 43 citations
Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo et al.ICCV 2019 · 1,125 citations
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 956 citations
Related papers
- Learning Deep Neural Networks under Agnostic Corrupted SupervisionBoyang Liu, Mengying Sun, Ding Wang, Pang-Ning Tan et al.ICML 2021 · 7 citations
- Tackling Instance-Dependent Label Noise via a Universal Probabilistic ModelQizhou Wang, Bo Han, Tongliang Liu, Gang Niu et al.AAAI 2021 · 34 citations
- Learning Discriminative Dynamics with Label Corruption for Noisy Label DetectionSuyeon Kim, Dongha Lee, SeongKu Kang, Sukang Chae et al.CVPR 2024
- A Topological Filter for Learning with Label NoisePengxiang Wu, Songzhu Zheng, Mayank Goswami, Dimitris N. Metaxas et al.NeurIPS 2020 · 143 citations
- Error-Bounded Correction of Noisy LabelsSongzhu Zheng, Pengxiang Wu, Aman Goswami, Mayank Goswami et al.ICML 2020 · 153 citations
