DeepGini: prioritizing massive tests to enhance the robustness of deep neural networks
Yang Feng, Qingkai Shi, Xinyu Gao, Jun Wan, Chunrong Fang, Zhenyu Chen
Abstract
Deep neural networks (DNN) have been deployed in many software systems to assist in various classification tasks. In company with the fantastic effectiveness in classification, DNNs could also exhibit incorrect behaviors and result in accidents and losses. Therefore, testing techniques that can detect incorrect DNN behaviors and improve DNN quality are extremely necessary and critical. However, the testing oracle, which defines the correct output for a given input, is often not available in the automated testing. To obtain the oracle information, the testing tasks of DNN-based systems usually require expensive human efforts to label the testing data, which significantly slows down the process of quality assurance. To mitigate this problem, we propose DeepGini, a test prioritization technique designed based on a statistical perspective of DNN. Such a statistical perspective allows us to reduce the problem of measuring misclassification probability to the problem of measuring set impurity, which allows us to quickly identify possibly-misclassified tests. To evaluate, we conduct an extensive empirical study on popular datasets and prevalent DNN models. The experimental results demonstrate that DeepGini outperforms existing coverage-based techniques in prioritizing tests regarding both effectiveness and efficiency. Meanwhile, we observe that the tests prioritized at the front by DeepGini are more effective in improving the DNN quality in comparison with the coverage-based techniques. CCS CONCEPTS • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers37
- Prioritizing Test Inputs for Deep Neural Networks via Mutation AnalysisZan Wang, Hanmo You, Junjie Chen, Yingyi Zhang et al.ICSE 2021 · 117 citations
- Copy, Right? A Testing Framework for Copyright Protection of Deep Learning ModelsJialuo Chen, Jingyi Wang, Tinglan Peng, Youcheng Sun et al.S&P 2022 · 94 citations
- Testing of autonomous driving systems: where are we and where should we go?Guannan Lou, Yao Deng, Xi Zheng, Mengshi Zhang et al.FSE 2022 · 85 citations
- DeepPayload: Black-box Backdoor Attack on Deep Learning Models through Neural Payload InjectionYuanchun Li, Jiayi Hua, Haoyu Wang, Chunyang Chen et al.ICSE 2021 · 70 citations
- RobOT: Robustness-Oriented Testing for Deep Learning SystemsJingyi Wang, Jialuo Chen, Youcheng Sun, Xingjun Ma et al.ICSE 2021 · 62 citations
Builds on1
Related papers
- Simple techniques work surprisingly well for neural network test prioritization and active learning (replicability study)Michael Weiss, Paolo TonellaISSTA 2022 · 50 citations
- Distribution-Aware Testing of Neural Networks Using Generative ModelsSwaroopa Dola, Matthew B. Dwyer, Mary Lou SoffaICSE 2021 · 3 citations
- Operation is the hardest teacher: estimating DNN accuracy looking for mispredictionsAntonio Guerriero, Roberto Pietrantuono, Stefano RussoICSE 2021 · 3 citations
- In Defense of Simple Techniques for Neural Network Test Case SelectionShenglin Bao, Chaofeng Sha, Bihuan Chen, Xin Peng et al.ISSTA 2023 · 9 citations
- DeepSample: DNN sampling-based testing for operational accuracy assessmentAntonio Guerriero, Roberto Pietrantuono, Stefano RussoICSE 2024 · 6 citations
