Detecting Corrupted Labels Without Training a Model to Predict
Zhaowei Zhu, Zihao Dong, Yang Liu
摘要
Label noise in real-world datasets encodes wrong correlation patterns and impairs the generalization of deep neural networks (DNNs). It is critical to find efficient ways to detect corrupted patterns. Current methods primarily focus on designing robust training techniques to prevent DNNs from memorizing corrupted patterns. These approaches often require customized training processes and may overfit corrupted patterns, leading to a performance drop in detection. In this paper, from a more data-centric perspective, we propose a training-free solution to detect corrupted labels. Intuitively, closer'' instances are more likely to share the same clean label. Based on the neighborhood information, we propose two methods: the first one uses local voting"via checking the noisy label consensuses of nearby features. The second one is a ranking-based approach that scores each instance and filters out a guaranteed number of instances that are likely to be corrupted. We theoretically analyze how the quality of features affects the local voting and provide guidelines for tuning neighborhood size. We also prove the worst-case error bound for the ranking-based method. Experiments with both synthetic and real-world label noise demonstrate our training-free solutions consistently and significantly improve most of the training-based baselines. Code is available at github.com/UCSC-REAL/SimiFeat.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Mitigating Neural Network Overconfidence with Logit NormalizationHongxin Wei, Renchunzi Xie, Hao Cheng, Lei Feng 等ICML 2022 · 被引用 386 次
- Mitigating Memorization of Noisy Labels by Clipping the Model PredictionHongxin Wei, Huiping Zhuang, Renchunzi Xie, Lei Feng 等ICML 2023 · 被引用 54 次
- Open-Sampling: Exploring Out-of-Distribution data for Re-balancing Long-tailed datasetsHongxin Wei, Lue Tao, Renchunzi Xie, Lei Feng 等ICML 2022 · 被引用 46 次
- The Rich Get Richer: Disparate Impact of Semi-Supervised LearningZhaowei Zhu, Tianyi Luo, Yang LiuICLR 2022 · 被引用 44 次
- Beyond Images: Label Noise Transition Matrix Estimation for Tasks with Lower-Quality FeaturesZhaowei Zhu, Jialu Wang, Yang LiuICML 2022 · 被引用 43 次
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 被引用 1,326 次
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo 等ICCV 2019 · 被引用 1,125 次
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 被引用 956 次
相关 Paper
- Learning Deep Neural Networks under Agnostic Corrupted SupervisionBoyang Liu, Mengying Sun, Ding Wang, Pang-Ning Tan 等ICML 2021 · 被引用 7 次
- Tackling Instance-Dependent Label Noise via a Universal Probabilistic ModelQizhou Wang, Bo Han, Tongliang Liu, Gang Niu 等AAAI 2021 · 被引用 34 次
- Learning Discriminative Dynamics with Label Corruption for Noisy Label DetectionSuyeon Kim, Dongha Lee, SeongKu Kang, Sukang Chae 等CVPR 2024
- A Topological Filter for Learning with Label NoisePengxiang Wu, Songzhu Zheng, Mayank Goswami, Dimitris N. Metaxas 等NeurIPS 2020 · 被引用 143 次
- Error-Bounded Correction of Noisy LabelsSongzhu Zheng, Pengxiang Wu, Aman Goswami, Mayank Goswami 等ICML 2020 · 被引用 153 次
