MisDetect: Iterative Mislabel Detection using Early Loss
Yuhao Deng, Chengliang Chai, Lei Cao, Nan Tang, Jiayi Wang, Ju Fan, Ye Yuan, Guoren Wang
Abstract
Supervised machine learning (ML) models trained on data with mislabeled instances often produce inaccurate results due to label errors. Traditional methods of detecting mislabeled instances rely on data proximity, where an instance is considered mislabeled if its label is inconsistent with its neighbors. However, it often performs poorly, because an instance does not always share the same label with its neighbors. ML-based methods instead utilize trained models to differentiate between mislabeled and clean instances. However, these methods struggle to achieve high accuracy, since the models may have already overfitted mislabeled instances.
In this paper, we propose a novel framework, MisDetect, that detects mislabeled instances during model training. MisDetect leverages the early loss observation to iteratively identify and remove mislabeled instances. In this process, influence-based verification is applied to enhance the detection accuracy. Moreover, MisDetect automatically determines when the early loss is no longer effective in detecting mislabels such that the iterative detection process should terminate. Finally, for the training instances that MisDetect is still not certain about whether they are mislabeled or not, MisDetect automatically produces some pseudo labels to learn a binary classification model and leverages the generalization ability of the machine learning model to determine their status. Our experiments on 15 datasets show that MisDetect outperforms 10 baseline methods, demonstrating its effectiveness in detecting mislabeled instances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b52f225a-5a1a-41b5-a443-0883deb85b02Cited by top-tier papers5
- Outlier Summarization via Human Interpretable RulesYuhao Deng, Yu Wang, Lei Cao, Lianpeng Qiao et al.VLDB 2024 · 6 citations
- Enhancing Sample Selection Against Label Noise by Cutting Mislabeled Easy ExamplesSuqin Yuan, Lei Feng, Bo Han, Tongliang LiuNeurIPS 2025 · 5 citations
- TuneAhead: Predicting Fine-tuning Performance Before Training BeginsYuxiang Luo, Haonan Long, Chen Wang, Qiqi Duan et al.ICML 2026
- Harnessing Diversity for Important Data Selection in Pretraining Large Language ModelsChi Zhang, Huaping Zhong, Kuan Zhang, Chengliang Chai et al.ICLR 2025
- A Risk Decomposition Framework for Pre-hoc Fine-tuning PredictionYuxiang Luo, Chen Wang, Nan TangICML 2026
Builds on13
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 633 citations
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang et al.ICDE 2021 · 127 citations
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li et al.VLDB 2022 · 62 citations
- Human-in-the-loop Outlier DetectionChengliang Chai, Lei Cao, Guoliang Li, Jian Li et al.SIGMOD 2020 · 57 citations
Related papers
- Learning Discriminative Dynamics with Label Corruption for Noisy Label DetectionSuyeon Kim, Dongha Lee, SeongKu Kang, Sukang Chae et al.CVPR 2024
- Deep k-NN for Noisy LabelsDara Bahri, Heinrich Jiang, Maya R. GuptaICML 2020 · 90 citations
- Learning from Training Dynamics: Identifying Mislabeled Data beyond Manually Designed FeaturesQingrui Jia, Xuhong Li, Lei Yu, Jiang Bian et al.AAAI 2023 · 12 citations
- Data Glitches Discovery using Influence-based Model ExplanationsNikolaos Myrtakis, Ioannis Tsamardinos, Vassilis ChristophidesKDD 2025
- Early Stopping Against Label Noise Without Validation DataSuqin Yuan, Lei Feng, Tongliang LiuICLR 2024 · 39 citations
