Learning from Training Dynamics: Identifying Mislabeled Data beyond Manually Designed Features
Qingrui Jia, Xuhong Li, Lei Yu, Jiang Bian, Penghao Zhao, Shupeng Li, Haoyi Xiong, Dejing Dou
摘要
While mislabeled or ambiguously-labeled samples in the training set could negatively affect the performance of deep models, diagnosing the dataset and identifying mislabeled samples helps to improve the generalization power. Training dynamics, i.e., the traces left by iterations of optimization algorithms, have recently been proved to be effective to localize mislabeled samples with hand-crafted features. In this paper, beyond manually designed features, we introduce a novel learning-based solution, leveraging a noise detector, instanced by an LSTM network, which learns to predict whether a sample was mislabeled using the raw training dynamics as input. Specifically, the proposed method trains the noise detector in a supervised manner using the dataset with synthesized label noises and can adapt to various datasets (either naturally or synthesized label-noised) without retraining. We conduct extensive experiments to evaluate the proposed method. We train the noise detector based on the synthesized label-noised CIFAR dataset and test such noise detector on Tiny ImageNet, CUB-200, Caltech-256, WebVision and Clothing1M. Results show that the proposed method precisely detects mislabeled samples on various datasets without further adaptation, and outperforms state-of-the-art methods. Besides, more experiments demonstrate that the mislabel identification can guide a label correction, namely data debugging, providing orthogonal improvements of algorithm-centric state-of-the-art techniques from the data aspect.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Dissecting Sample Hardness: A Fine-Grained Analysis of Hardness Characterization Methods for Data-Centric AINabeel Seedat, Fergus Imrie, Mihaela van der SchaarICLR 2024 · 被引用 16 次
- SAP: Corrective Machine Unlearning with Scaled Activation Projection for Label Noise RobustnessSangamesh Kodge, Deepak Ravikumar, Gobinda Saha, Kaushik RoyAAAI 2025 · 被引用 10 次
- Enhanced Sample Selection with Confidence Tracking: Identifying Correctly Labeled Yet Hard-to-Learn Samples in Noisy DataWeiran Pan, Wei Wei, Feida Zhu, Yong DengAAAI 2025 · 被引用 6 次
- Debiased Sample Selection for Learning with Noisy LabelsWeiran Pan, Wei Wei, Wenfeng XieCVPR 2026
- Non-Stationary Predictions May Be More Informative: Exploring Pseudo-Labels with a Two-Phase Pattern of Training DynamicsHongbin Pei, Jingxin Hai, Yu Li, Huiqi Deng 等ICML 2025
它引用的顶会 Paper9
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 被引用 1,326 次
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo 等ICCV 2019 · 被引用 1,125 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- SELF: Learning to Filter Noisy Labels with Self-EnsemblingDuc Tam Nguyen, Chaithanya Kumar Mummadi, Thi-Phuong-Nhung Ngo, Thi Hoai Phuong Nguyen 等ICLR 2020 · 被引用 354 次
- Robust early-learning: Hindering the memorization of noisy labelsXiaobo Xia, Tongliang Liu, Bo Han, Chen Gong 等ICLR 2021 · 被引用 322 次
相关 Paper
- Learning Discriminative Dynamics with Label Corruption for Noisy Label DetectionSuyeon Kim, Dongha Lee, SeongKu Kang, Sukang Chae 等CVPR 2024
- Detecting Corrupted Labels Without Training a Model to PredictZhaowei Zhu, Zihao Dong, Yang LiuICML 2022 · 被引用 84 次
- MisDetect: Iterative Mislabel Detection using Early LossYuhao Deng, Chengliang Chai, Lei Cao, Nan Tang 等VLDB 2024 · 被引用 13 次
- Noise Attention Learning: Enhancing Noise Robustness by Gradient ScalingYangdi Lu, Yang Bo, Wenbo HeNeurIPS 2022 · 被引用 13 次
- Early Stopping Against Label Noise Without Validation DataSuqin Yuan, Lei Feng, Tongliang LiuICLR 2024 · 被引用 39 次
