On Revisiting Entropy for Identifying Mislabeled Images
Chunlei Li, Zixuan Zheng, Yilei Shi, Guanglu Dong, Pengfei Li, Jingliang Hu, Xiao Zhu, Lichao Mou
摘要
Mislabeled samples in training datasets severely degrade the performance of deep networks, as overparameterized models tend to memorize erroneous labels. We address this challenge by proposing a novel approach for mislabeled data detection that leverages training dynamics. Our method is grounded in the key observation that correctly labeled samples exhibit consistent entropy decrease during training, while mislabeled samples maintain relatively high entropy throughout the training process. Building on this insight, we introduce a signed entropy integral (SEI) statistic that captures both the magnitude and temporal trend of prediction entropy across training epochs. SEI is broadly applicable to classification networks and demonstrates particular effectiveness when integrated with contrastive language-image pretraining (CLIP) architectures. Through extensive experiments on four medical imaging datasets---a domain particularly susceptible to labeling errors due to diagnostic complexity---spanning diverse modalities and pathologies, we demonstrate that SEI achieves state-of-the-art performance in mislabeled data identification, outperforming existing methods while maintaining computational efficiency and implementation simplicity. Our code is available at https://github.com/MedAITech/SEI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo 等ICCV 2019 · 被引用 1,125 次
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 被引用 798 次
- Semi-supervised Medical Image Segmentation through Dual-task ConsistencyXiangde Luo, Jieneng Chen, Tao Song, Guotai WangAAAI 2021 · 被引用 754 次
相关 Paper
- Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical AnalysisHanbin Ko, Chang-Min ParkCVPR 2025
- EncoderMI: Membership Inference against Pre-trained Encoders in Contrastive LearningHongbin Liu, Jinyuan Jia, Wenjie Qu, Neil Zhenqiang GongCCS 2021 · 被引用 63 次
- CLIP-Driven Universal Model for Organ Segmentation and Tumor DetectionJie Liu, Yixiao Zhang, Jieneng Chen, Junfei Xiao 等ICCV 2023 · 被引用 336 次
- Detecting Backdoor Samples in Contrastive Language Image PretrainingHanxun Huang, Sarah Monazam Erfani, Yige Li, Xingjun Ma 等ICLR 2025
- Federated CLIP for Resource-Efficient Heterogeneous Medical Image ClassificationYihang Wu, Ahmad ChaddadAAAI 2026 · 被引用 1 次
