DeepMetis: Augmenting a Deep Learning Test Set to Increase its Mutation Score
Vincenzo Riccio, Nargiz Humbatova, Gunel Jahangirova, Paolo Tonella
Abstract
Deep Learning (DL) components are routinely integrated into software systems that need to perform complex tasks such as image or natural language processing. The adequacy of the test data used to test such systems can be assessed by their ability to expose artificially injected faults (mutations) that simulate real DL faults. In this paper, we describe an approach to automatically generate new test inputs that can be used to augment the existing test set so that its capability to detect DL mutations increases. Our tool DEEPMETIS implements a search based input generation strategy. To account for the non-determinism of the training and the mutation processes, our fitness function involves multiple instances of the DL model under test. Experimental results show that DEEPMETIS is effective at augmenting the given test set, increasing its capability to detect mutants by 63% on average. A leave-one-out experiment shows that the augmented test set is capable of exposing unseen mutants, which simulate the occurrence of yet undetected faults.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2377fb98-1a34-41d3-b18a-981b060f4d50Cited by top-tier papers4
- When and Why Test Generators for Deep Learning Produce Invalid Inputs: an Empirical StudyVincenzo Riccio, Paolo TonellaICSE 2023 · 29 citations
- Decomposition of Deep Neural Networks into Modules via Mutation AnalysisAli GhanbariISSTA 2024 · 4 citations
- Using Fourier Analysis and Mutant Clustering to Accelerate DNN Mutation TestingAli Ghanbari, Sasan TavakkolASE 2025 · 1 citation
- AudioTest: Prioritizing Audio Test CasesYinghua Li, Xueqi Dang, Wendkûuni C. Ouédraogo, Jacques Klein et al.ISSTA 2025 · 1 citation
Builds on8
- Taxonomy of real faults in deep learning systemsNargiz Humbatova, Gunel Jahangirova, Gabriele Bavota, Vincenzo Riccio et al.ICSE 2020 · 281 citations
- Is neuron coverage a meaningful measure for testing deep neural networks?Fabrice Harel-Canada, Lingxiao Wang, Muhammad Ali Gulzar, Quanquan Gu et al.FSE 2020 · 149 citations
- Model-based exploration of the frontier of behaviours for deep learning system testingVincenzo Riccio, Paolo TonellaFSE 2020 · 134 citations
- DeepCrime: mutation testing of deep learning systems based on real faultsNargiz Humbatova, Gunel Jahangirova, Paolo TonellaISSTA 2021 · 114 citations
- DeepHyperion: exploring the feature space of deep learning-based systems through illumination searchTahereh Zohdinasab, Vincenzo Riccio, Alessio Gambi, Paolo TonellaISSTA 2021 · 76 citations
Related papers
- Fuzz testing based data augmentation to improve robustness of deep neural networksXiang Gao, Ripon K. Saha, Mukul R. Prasad, Abhik RoychoudhuryICSE 2020 · 116 citations
- Dynamic Data Fault Localization for Deep Neural NetworksYining Yin, Yang Feng, Shihao Weng, Zixi Liu et al.FSE 2023 · 10 citations
- Repairing Failure-inducing Inputs with Input ReflectionYan Xiao, Yun Lin, Ivan Beschastnikh, Changsheng Sun et al.ASE 2022 · 8 citations
- Regression Fuzzing for Deep Learning SystemsHanmo You, Zan Wang, Junjie Chen, Shuang Liu et al.ICSE 2023 · 28 citations
- Graph-based Fuzz Testing for Deep Learning Inference EnginesWeisi Luo, Dong Chai, Xiaoyue Run, Jiang Wang et al.ICSE 2021 · 37 citations
