MisDetect: Iterative Mislabel Detection using Early Loss
Yuhao Deng, Chengliang Chai, Lei Cao, Nan Tang, Jiayi Wang, Ju Fan, Ye Yuan, Guoren Wang
摘要
Supervised machine learning (ML) models trained on data with mislabeled instances often produce inaccurate results due to label errors. Traditional methods of detecting mislabeled instances rely on data proximity, where an instance is considered mislabeled if its label is inconsistent with its neighbors. However, it often performs poorly, because an instance does not always share the same label with its neighbors. ML-based methods instead utilize trained models to differentiate between mislabeled and clean instances. However, these methods struggle to achieve high accuracy, since the models may have already overfitted mislabeled instances.
In this paper, we propose a novel framework, MisDetect, that detects mislabeled instances during model training. MisDetect leverages the early loss observation to iteratively identify and remove mislabeled instances. In this process, influence-based verification is applied to enhance the detection accuracy. Moreover, MisDetect automatically determines when the early loss is no longer effective in detecting mislabels such that the iterative detection process should terminate. Finally, for the training instances that MisDetect is still not certain about whether they are mislabeled or not, MisDetect automatically produces some pseudo labels to learn a binary classification model and leverages the generalization ability of the machine learning model to determine their status. Our experiments on 15 datasets show that MisDetect outperforms 10 baseline methods, demonstrating its effectiveness in detecting mislabeled instances.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Outlier Summarization via Human Interpretable RulesYuhao Deng, Yu Wang, Lei Cao, Lianpeng Qiao 等VLDB 2024 · 被引用 6 次
- Enhancing Sample Selection Against Label Noise by Cutting Mislabeled Easy ExamplesSuqin Yuan, Lei Feng, Bo Han, Tongliang LiuNeurIPS 2025 · 被引用 5 次
- TuneAhead: Predicting Fine-tuning Performance Before Training BeginsYuxiang Luo, Haonan Long, Chen Wang, Qiqi Duan 等ICML 2026
- Harnessing Diversity for Important Data Selection in Pretraining Large Language ModelsChi Zhang, Huaping Zhong, Kuan Zhang, Chengliang Chai 等ICLR 2025
- A Risk Decomposition Framework for Pre-hoc Fine-tuning PredictionYuxiang Luo, Chen Wang, Nan TangICML 2026
它引用的顶会 Paper13
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 被引用 1,326 次
- Certified Data Removal from Machine Learning ModelsChuan Guo, Tom Goldstein, Awni Y. Hannun, Laurens van der MaatenICML 2020 · 被引用 633 次
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang 等ICDE 2021 · 被引用 127 次
- Selective Data Acquisition in the Wild for Model ChargingChengliang Chai, Jiabin Liu, Nan Tang, Guoliang Li 等VLDB 2022 · 被引用 62 次
- Human-in-the-loop Outlier DetectionChengliang Chai, Lei Cao, Guoliang Li, Jian Li 等SIGMOD 2020 · 被引用 57 次
相关 Paper
- Learning Discriminative Dynamics with Label Corruption for Noisy Label DetectionSuyeon Kim, Dongha Lee, SeongKu Kang, Sukang Chae 等CVPR 2024
- Deep k-NN for Noisy LabelsDara Bahri, Heinrich Jiang, Maya R. GuptaICML 2020 · 被引用 90 次
- Learning from Training Dynamics: Identifying Mislabeled Data beyond Manually Designed FeaturesQingrui Jia, Xuhong Li, Lei Yu, Jiang Bian 等AAAI 2023 · 被引用 12 次
- Data Glitches Discovery using Influence-based Model ExplanationsNikolaos Myrtakis, Ioannis Tsamardinos, Vassilis ChristophidesKDD 2025
- Early Stopping Against Label Noise Without Validation DataSuqin Yuan, Lei Feng, Tongliang LiuICLR 2024 · 被引用 39 次
