On Revisiting Entropy for Identifying Mislabeled Images
Chunlei Li, Zixuan Zheng, Yilei Shi, Guanglu Dong, Pengfei Li, Jingliang Hu, Xiao Zhu, Lichao Mou
Abstract
Mislabeled samples in training datasets severely degrade the performance of deep networks, as overparameterized models tend to memorize erroneous labels. We address this challenge by proposing a novel approach for mislabeled data detection that leverages training dynamics. Our method is grounded in the key observation that correctly labeled samples exhibit consistent entropy decrease during training, while mislabeled samples maintain relatively high entropy throughout the training process. Building on this insight, we introduce a signed entropy integral (SEI) statistic that captures both the magnitude and temporal trend of prediction entropy across training epochs. SEI is broadly applicable to classification networks and demonstrates particular effectiveness when integrated with contrastive language-image pretraining (CLIP) architectures. Through extensive experiments on four medical imaging datasets---a domain particularly susceptible to labeling errors due to diagnostic complexity---spanning diverse modalities and pathologies, we demonstrate that SEI achieves state-of-the-art performance in mislabeled data identification, outperforming existing methods while maintaining computational efficiency and implementation simplicity. Our code is available at https://github.com/MedAITech/SEI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo et al.ICCV 2019 · 1,125 citations
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 798 citations
- Semi-supervised Medical Image Segmentation through Dual-task ConsistencyXiangde Luo, Jieneng Chen, Tao Song, Guotai WangAAAI 2021 · 754 citations
Related papers
- Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical AnalysisHanbin Ko, Chang-Min ParkCVPR 2025
- EncoderMI: Membership Inference against Pre-trained Encoders in Contrastive LearningHongbin Liu, Jinyuan Jia, Wenjie Qu, Neil Zhenqiang GongCCS 2021 · 63 citations
- CLIP-Driven Universal Model for Organ Segmentation and Tumor DetectionJie Liu, Yixiao Zhang, Jieneng Chen, Junfei Xiao et al.ICCV 2023 · 336 citations
- Detecting Backdoor Samples in Contrastive Language Image PretrainingHanxun Huang, Sarah Monazam Erfani, Yige Li, Xingjun Ma et al.ICLR 2025
- Federated CLIP for Resource-Efficient Heterogeneous Medical Image ClassificationYihang Wu, Ahmad ChaddadAAAI 2026 · 1 citation
