Enhanced Sample Selection with Confidence Tracking: Identifying Correctly Labeled Yet Hard-to-Learn Samples in Noisy Data
Weiran Pan, Wei Wei, Feida Zhu, Yong Deng
摘要
We propose a novel sample selection method for image classification in the presence of noisy labels. Existing methods typically consider small-loss samples as correctly labeled. However, some correctly labeled samples are inherently difficult for the model to learn and can exhibit high loss similar to mislabeled samples in the early stages of training. Consequently, setting a threshold on per-sample loss to select correct labels results in a trade-off between precision and recall in sample selection: a lower threshold may miss many correctly labeled hard-to-learn samples (low recall), while a higher threshold may include many mislabeled samples (low precision). To address this issue, our goal is to accurately distinguish correctly labeled yet hard-to-learn samples from mislabeled ones, thus alleviating the trade-off dilemma. We achieve this by considering the trends in model prediction confidence rather than relying solely on loss values. Empirical observations show that only for correctly labeled samples, the model's prediction confidence for the annotated labels typically increases faster than for any other classes. Based on this insight, we propose tracking the confidence gaps between the annotated labels and other classes during training and evaluating their trends using the Mann-Kendall Test. A sample is considered potentially correctly labeled if all its confidence gaps tend to increase. Our method functions as a plug-and-play component that can be seamlessly integrated into existing sample selection techniques. Experiments on several standard benchmarks and real-world datasets demonstrate that our method enhances the performance of existing methods for learning with noisy labels.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Debiased Sample Selection for Learning with Noisy LabelsWeiran Pan, Wei Wei, Wenfeng XieCVPR 2026
- Identifying and Correcting Label Noise for Robust GNNs via Influence ContradictionWei Ju, Wei Zhang, Siyu Yi, Zhengyang Mao 等ICML 2026
- LANE: Label-Aware Noise Elimination for Fine-Grained Text ClassificationTiberiu Sosea, Cornelia CarageaICLR 2026
它引用的顶会 Paper35
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 被引用 1,326 次
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo 等ICCV 2019 · 被引用 1,125 次
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 被引用 798 次
- Normalized Loss Functions for Deep Learning with Noisy LabelsXingjun Ma, Hanxun Huang, Yisen Wang, Simone Romano 等ICML 2020 · 被引用 547 次
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 被引用 398 次
相关 Paper
- Enhancing Sample Selection Against Label Noise by Cutting Mislabeled Easy ExamplesSuqin Yuan, Lei Feng, Bo Han, Tongliang LiuNeurIPS 2025 · 被引用 5 次
- Sample Selection with Uncertainty of Losses for Learning with Noisy LabelsXiaobo Xia, Tongliang Liu, Bo Han, Mingming Gong 等ICLR 2022 · 被引用 139 次
- RankMatch: Fostering Confidence and Consistency in Learning with Noisy LabelsZiyi Zhang, Weikai Chen, Chaowei Fang, Zhen Li 等ICCV 2023 · 被引用 11 次
- Sample-wise Label Confidence Incorporation for Learning with Noisy LabelsChanho Ahn, Kikyung Kim, Ji-Won Baek, Jongin Lim 等ICCV 2023 · 被引用 11 次
- Confidence-based Reliable Learning under Dual NoisesPeng Cui, Yang Yue, Zhijie Deng, Jun ZhuNeurIPS 2022 · 被引用 13 次
