Evaluating Deep Neural Networks in Deployment: A Comparative Study (Replicability Study)
Eduard Pinconschi, Divya Gopinath, Rui Abreu, Corina S. Pasareanu
Abstract
As deep neural networks (DNNs) are increasingly used in safety-critical applications, there is a growing concern for their reliability. Even highly trained, high-performant networks are not 100% accurate. However, it is very difficult to predict their behavior during deployment without ground truth. In this paper, we provide a comparative and replicability study on recent approaches that have been proposed to evaluate the reliability of DNNs in deployment. We find that it is hard to run and reproduce the results for these approaches on their replication packages and even more difficult to run them on artifacts other than their own. Further, it is difficult to compare the effectiveness of the approaches, due to the lack of clearly defined evaluation metrics. Our results indicate that more effort is needed in our research community to obtain sound techniques for evaluating the reliability of neural networks in safety-critical domains. To this end, we contribute an evaluation framework that incorporates the considered approaches and enables evaluation on common benchmarks, using common metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8bf13dc1-2ba9-4c6a-aaea-f796e6bfec33Builds on5
- Misbehaviour prediction for autonomous driving systemsAndrea Stocco, Michael Weiss, Marco Calzana, Paolo TonellaICSE 2020 · 138 citations
- Dissector: input validation for deep learning applications by crossing-layer dissectionHuiyan Wang, Jingwei Xu, Chang Xu, Xiaoxing Ma et al.ICSE 2020 · 56 citations
- Self-Checking Deep Neural Networks in DeploymentYan Xiao, Ivan Beschastnikh, David S. Rosenblum, Changsheng Sun et al.ICSE 2021 · 36 citations
- Inferring Data Preconditions from Deep Learning Models for Trustworthy Prediction in DeploymentShibbir Ahmed, Hongyang Gao, Hridesh RajanICSE 2024 · 3 citations
- A Programmatic and Semantic Approach to Explaining and Debugging Neural Network Based Object DetectorsEdward Kim, Divya Gopinath, Corina S. Pasareanu, Sanjit A. SeshiaCVPR 2020
Related papers
- Distribution-Aware Testing of Neural Networks Using Generative ModelsSwaroopa Dola, Matthew B. Dwyer, Mary Lou SoffaICSE 2021 · 3 citations
- Synthesizing Boxes Preconditions for Deep Neural NetworksZengyu Liu, Liqian Chen, Wanwei Liu, Ji WangISSTA 2024
- Reliability Assurance for Deep Neural Network Architectures Against Numerical DefectsLinyi Li, Yuhao Zhang, Luyao Ren, Yingfei Xiong et al.ICSE 2023 · 7 citations
- Reducing DNN Properties to Enable Falsification with Adversarial AttacksDavid Shriver, Sebastian G. Elbaum, Matthew B. DwyerICSE 2021 · 18 citations
- Test Selection for Deep Neural Networks using Meta-Models with Uncertainty MetricsDemet Demir, Aysu Betin Can, Elif SürerISSTA 2024 · 3 citations
