Robust Learning of Deep Predictive Models from Noisy and Imbalanced Software Engineering Datasets
Zhong Li, Minxue Pan, Yu Pei, Tian Zhang, Linzhang Wang, Xuandong Li
Abstract
With the rapid development of Deep Learning, deep predictive models have been widely applied to improve Software Engineering tasks, such as defect prediction and issue classification, and have achieved remarkable success. They are mostly trained in a supervised manner, which heavily relies on high-quality datasets. Unfortunately, due to the nature and source of software engineering data, the real-world datasets often suffer from the issues of sample mislabelling and class imbalance, thus undermining the effectiveness of deep predictive models in practice. This problem has become a major obstacle for deep learning-based Software Engineering. In this paper, we propose RobustTrainer, the first approach to learning deep predictive models on raw training datasets where the mislabelled samples and the imbalanced classes coexist. Robust-Trainer consists of a two-stage training scheme, where the first learns feature representations robust to sample mislabelling and the second builds a classifier robust to class imbalance based on the learned representations in the first stage. We apply RobustTrainer to two popular Software Engineering tasks, i.e., Bug Report Classification and Software Defect Prediction. Evaluation results show that RobustTrainer effectively tackles the mislabelling and class imbalance issues and produces significantly better deep predictive models compared to the other six comparison approaches. CCS CONCEPTS • Software and its engineering → Software development techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- DistXplore: Distribution-Guided Testing for Evaluating and Enhancing Deep Learning SystemsLongtian Wang, Xiaofei Xie, Xiaoning Du, Meng Tian et al.FSE 2023 · 15 citations
- An Empirical Study on Noisy Label Learning for Program UnderstandingWenhan Wang, Yanzhou Li, Anran Li, Jian Zhang et al.ICSE 2024 · 5 citations
Builds on10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain et al.ICLR 2021 · 937 citations
- Deep Self-Learning From Noisy LabelsJiangfan Han, Ping Luo, Xiaogang WangICCV 2019 · 315 citations
Related papers
- On Distribution Shift in Learning-based Bug DetectorsJingxuan He, Luca Beurer-Kellner, Martin T. VechevICML 2022 · 20 citations
- Does data sampling improve deep learning-based vulnerability detection? Yeas! and Nays!Xu Yang, Shaowei Wang, Yi Li, Shaohua WangICSE 2023 · 19 citations
- Explaining mispredictions of machine learning models using rule inductionJürgen Cito, Isil Dillig, Seohyun Kim, Vijayaraghavan Murali et al.FSE 2021 · 26 citations
- A Causal Learning Framework for Enhancing Robustness of Source Code ModelsJunyao Ye, Zhen Li, Xi Tang, Deqing Zou et al.FSE 2025
- AUTOTRAINER: An Automatic DNN Training Problem Detection and Repair SystemXiaoyu Zhang, Juan Zhai, Shiqing Ma, Chao ShenICSE 2021 · 62 citations
