Unveiling Hidden DNN Defects with Decision-Based Metamorphic Testing
Yuanyuan Yuan, Qi Pang, Shuai Wang
Abstract
Contemporary DNN testing works are frequently conducted using metamorphic testing (MT). In general, de facto MT frameworks mutate DNN input images using semantics-preserving mutations and determine if DNNs can yield consistent predictions. Nevertheless, we find that DNNs may rely on erroneous decisions (certain components on the DNN inputs) to make predictions, which may still retain the outputs by chance. Such DNN defects would be neglected by existing MT frameworks. Erroneous decisions, however, would likely result in successive mis-predictions over diverse images that may exist in real-life scenarios. This research aims to unveil the pervasiveness of hidden DNN defects caused by incorrect DNN decisions (but retaining consistent DNN predictions). To do so, we tailor and optimize modern eXplainable AI (XAI) techniques to identify visual concepts that represent regions in an input image upon which the DNN makes predictions. Then, we extend existing MT-based DNN testing frameworks to check the consistency of DNN decisions made over a test input and its mutated inputs. Our evaluation shows that existing MT frameworks are oblivious to a considerable number of DNN defects caused by erroneous decisions. We conduct human evaluations to justify the validity of our findings and to elucidate their characteristics. Through the lens of DNN decision-based metamorphic relations, we re-examine the effectiveness of metamorphic transformations proposed by existing MT frameworks. We summarize lessons from this study, which can provide insights and guidelines for future DNN testing. CCS Concepts • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca872ef7-ab85-41c1-b863-0edd95df9ab8Cited by top-tier papers9
- Explain Any Concept: Segment Anything Meets Concept-Based ExplanationAo Sun, Pingchuan Ma, Yuanyuan Yuan, Shuai WangNeurIPS 2023 · 69 citations
- CCTEST: Testing and Repairing Code Completion SystemsZongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang et al.ICSE 2023 · 49 citations
- FedSlice: Protecting Federated Learning Models from Malicious Participants with Model SlicingZiqi Zhang, Yuanchun Li, Bingyan Liu, Yifeng Cai et al.ICSE 2023 · 8 citations
- BinAug: Enhancing Binary Similarity Analysis with Low-Cost Input RepairingWai Kin Wong, Huaijin Wang, Zongjie Li, Shuai WangICSE 2024 · 4 citations
- See the Forest, not Trees: Unveiling and Escaping the Pitfalls of Error-Triggering Inputs in Neural Network TestingYuanyuan Yuan, Shuai Wang, Zhendong SuISSTA 2024 · 1 citation
Builds on7
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Algorithmic Transparency via Quantitative Input Influence: Theory and Experiments with Learning SystemsAnupam Datta, Shayak Sen, Yair ZickS&P 2016 · 774 citations
- Metamorphic Object Insertion for Testing Object Detection SystemsShuai Wang, Zhendong SuASE 2020 · 69 citations
- CCTEST: Testing and Repairing Code Completion SystemsZongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang et al.ICSE 2023 · 49 citations
- MT-Teql: Evaluating and Augmenting Neural NLIDB on Real-world Linguistic and Schema VariationsPingchuan Ma, Shuai WangVLDB 2022 · 38 citations
Related papers
- ASRTest: automated testing for deep-neural-network-driven speech recognition systemsPin Ji, Yang Feng, Jia Liu, Zhihong Zhao et al.ISSTA 2022 · 22 citations
- DevMuT: Testing Deep Learning Framework via Developer Expertise-Based MutationYanzhou Mu, Juan Zhai, Chunrong Fang, Xiang Chen et al.ASE 2024 · 2 citations
- Improving Deep Learning Framework Testing with Model-Level Metamorphic TestingYanzhou Mu, Juan Zhai, Chunrong Fang, Xiang Chen et al.ISSTA 2025 · 1 citation
- DeepMetis: Augmenting a Deep Learning Test Set to Increase its Mutation ScoreVincenzo Riccio, Nargiz Humbatova, Gunel Jahangirova, Paolo TonellaASE 2021 · 41 citations
- Distribution-Aware Testing of Neural Networks Using Generative ModelsSwaroopa Dola, Matthew B. Dwyer, Mary Lou SoffaICSE 2021 · 3 citations
