Operation is the hardest teacher: estimating DNN accuracy looking for mispredictions
Antonio Guerriero, Roberto Pietrantuono, Stefano Russo
Abstract
Deep Neural Networks (DNN) are typically tested for accuracy relying on a set of unlabelled real world data (operational dataset), from which a subset is selected, manually labelled and used as test suite. This subset is required to be small (due to manual labelling cost) yet to faithfully represent the operational context, with the resulting test suite containing roughly the same proportion of examples causing misprediction (i.e., failing test cases) as the operational dataset. However, while testing to estimate accuracy, it is desirable to also learn as much as possible from the failing tests in the operational dataset, since they inform about possible bugs of the DNN. A smart sampling strategy may allow to intentionally include in the test suite many examples causing misprediction, thus providing this way more valuable inputs for DNN improvement while preserving the ability to get trustworthy unbiased estimates. This paper presents a test selection technique (DeepEST) that actively looks for failing test cases in the operational dataset of a DNN, with the goal of assessing the DNN expected accuracy by a small and "informative" test suite (namely with a high number of mispredictions) for subsequent DNN improvement. Experiments with five subjects, combining four DNN models and three datasets, are described. The results show that DeepEST provides DNN accuracy estimates with precision close to (and often better than) those of existing sampling-based DNN testing techniques, while detecting from 5 to 30 times more mispredictions, with the same test suite size.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de94ea34-3f8a-4618-9ed2-06221983b49dCited by top-tier papers5
- RISE: robust wireless sensing using probabilistic and statistical assessmentsShuangjiao Zhai, Zhanyong Tang, Petteri Nurmi, Dingyi Fang et al.MobiCom 2021 · 20 citations
- Reliability Assurance for Deep Neural Network Architectures Against Numerical DefectsLinyi Li, Yuhao Zhang, Luyao Ren, Yingfei Xiong et al.ICSE 2023 · 7 citations
- DeepSample: DNN sampling-based testing for operational accuracy assessmentAntonio Guerriero, Roberto Pietrantuono, Stefano RussoICSE 2024 · 6 citations
- Empirical Insights of Test Selection Metrics under Multiple Testing Objectives and Distribution ShiftsJingyu Zhang, Fan Wang, Jacky Keung, Yihan Liao et al.FSE 2026
- Test Case Selection for Deep Neural Networks: A Replication Study on LLMs for Code (Replicability Study)Ali Asgari, Mitchell Olsthoorn, Annibale PanichellaISSTA 2026
Related papers
- DeepGini: prioritizing massive tests to enhance the robustness of deep neural networksYang Feng, Qingkai Shi, Xinyu Gao, Jun Wan et al.ISSTA 2020 · 206 citations
- Adaptive Test Selection for Deep Neural NetworksXinyu Gao, Yang Feng, Yining Yin, Zixi Liu et al.ICSE 2022 · 53 citations
- Distance-Aware Test Input Selection for Deep Neural NetworksZhong Li, Zhengfeng Xu, Ruihua Ji, Minxue Pan et al.ISSTA 2024 · 4 citations
- Distribution-Aware Testing of Neural Networks Using Generative ModelsSwaroopa Dola, Matthew B. Dwyer, Mary Lou SoffaICSE 2021 · 3 citations
- Repairing Failure-inducing Inputs with Input ReflectionYan Xiao, Yun Lin, Ivan Beschastnikh, Changsheng Sun et al.ASE 2022 · 8 citations
