Dynamic Data Fault Localization for Deep Neural Networks
Yining Yin, Yang Feng, Shihao Weng, Zixi Liu, Yuan Yao, Yichi Zhang, Zhihong Zhao, Zhenyu Chen
Abstract
Rich datasets have empowered various deep learning (DL) applications, leading to remarkable success in many fields. However, data faults hidden in the datasets could result in DL applications behaving unpredictably and even cause massive monetary and life losses. To alleviate this problem, in this paper, we propose a dynamic data fault localization approach, namely DFauLo, to locate the mislabeled and noisy data in the deep learning datasets. DFauLo is inspired by the conventional mutation-based code fault localization, but utilizes the differences between DNN mutants to amplify and identify the potential data faults. Specifically, it first generates multiple DNN model mutants of the original trained model. Then it extracts features from these mutants and maps them into a suspiciousness score indicating the probability of the given data being a data fault. Moreover, DFauLo is the first dynamic data fault localization technique, prioritizing the suspected data based on user feedback, and providing the generalizability to unseen data faults during training. To validate DFauLo, we extensively evaluate it on 26 cases with various fault types, data types, and model structures. We also evaluate DFauLo on three widely-used benchmark datasets. The results show that DFauLo outperforms the state-of-the-art techniques in almost all cases and locates hundreds of different types of real data faults in benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8f2fb5c-81fa-47eb-aa56-32ad3c90b59aBuilds on22
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- NLNL: Negative Learning for Noisy LabelsYoungdong Kim, Junho Yim, Juseung Yun, Junmo KimICCV 2019 · 338 citations
- Taxonomy of real faults in deep learning systemsNargiz Humbatova, Gunel Jahangirova, Gabriele Bavota, Vincenzo Riccio et al.ICSE 2020 · 281 citations
- DeepGini: prioritizing massive tests to enhance the robustness of deep neural networksYang Feng, Qingkai Shi, Xinyu Gao, Jun Wan et al.ISSTA 2020 · 206 citations
Related papers
- Mutation-based Fault Localization of Deep Neural NetworksAli Ghanbari, Deepak-George Thomas, Muhammad Arbab Arshad, Hridesh RajanASE 2023 · 20 citations
- DeepLocalize: Fault Localization for Deep Neural NetworksMohammad Wardat, Wei Le, Hridesh RajanICSE 2021 · 93 citations
- DeepMetis: Augmenting a Deep Learning Test Set to Increase its Mutation ScoreVincenzo Riccio, Nargiz Humbatova, Gunel Jahangirova, Paolo TonellaASE 2021 · 41 citations
- DeepFD: Automated Fault Diagnosis and Localization for Deep Learning ProgramsJialun Cao, Meiziniu Li, Xiao Chen, Ming Wen et al.ICSE 2022 · 42 citations
- DeepCrime: mutation testing of deep learning systems based on real faultsNargiz Humbatova, Gunel Jahangirova, Paolo TonellaISSTA 2021 · 114 citations
