When and Why Test Generators for Deep Learning Produce Invalid Inputs: an Empirical Study
Vincenzo Riccio, Paolo Tonella
摘要
Testing Deep Learning (DL) based systems inherently requires large and representative test sets to evaluate whether DL systems generalise beyond their training datasets. Diverse Test Input Generators (TIGs) have been proposed to produce artificial inputs that expose issues of the DL systems by triggering misbehaviours. Unfortunately, such generated inputs may be invalid, i.e., not recognisable as part of the input domain, thus providing an unreliable quality assessment. Automated validators can ease the burden of manually checking the validity of inputs for human testers, although input validity is a concept difficult to formalise and, thus, automate. In this paper, we investigate to what extent TIGs can generate valid inputs, according to both automated and human validators. We conduct a large empirical study, involving 2 different automated validators, 220 human assessors, 5 different TIGs and 3 classification tasks. Our results show that 84% artificially generated inputs are valid, according to automated validators, but their expected label is not always preserved. Automated validators reach a good consensus with humans (78% accuracy), but still have limitations when dealing with feature-rich datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Bridging the Gap between Real-world and Synthetic Images for Testing Autonomous Driving SystemsMohammad Hossein Amini, Shiva NejatiASE 2024 · 被引用 2 次
- See the Forest, not Trees: Unveiling and Escaping the Pitfalls of Error-Triggering Inputs in Neural Network TestingYuanyuan Yuan, Shuai Wang, Zhendong SuISSTA 2024 · 被引用 1 次
- AudioTest: Prioritizing Audio Test CasesYinghua Li, Xueqi Dang, Wendkûuni C. Ouédraogo, Jacques Klein 等ISSTA 2025 · 被引用 1 次
- TestifAI: Tomography-Based Testing for Deep Learning SystemsArooj Arif, Tobias Hartung, Elena Botoeva, Alexandros KoliousisICSE 2026
它引用的顶会 Paper9
- Taxonomy of real faults in deep learning systemsNargiz Humbatova, Gunel Jahangirova, Gabriele Bavota, Vincenzo Riccio 等ICSE 2020 · 被引用 281 次
- Is neuron coverage a meaningful measure for testing deep neural networks?Fabrice Harel-Canada, Lingxiao Wang, Muhammad Ali Gulzar, Quanquan Gu 等FSE 2020 · 被引用 149 次
- Misbehaviour prediction for autonomous driving systemsAndrea Stocco, Michael Weiss, Marco Calzana, Paolo TonellaICSE 2020 · 被引用 138 次
- Model-based exploration of the frontier of behaviours for deep learning system testingVincenzo Riccio, Paolo TonellaFSE 2020 · 被引用 134 次
- DeepHyperion: exploring the feature space of deep learning-based systems through illumination searchTahereh Zohdinasab, Vincenzo Riccio, Alessio Gambi, Paolo TonellaISSTA 2021 · 被引用 76 次
相关 Paper
- Distribution-Aware Testing of Neural Networks Using Generative ModelsSwaroopa Dola, Matthew B. Dwyer, Mary Lou SoffaICSE 2021 · 被引用 3 次
- DeepMetis: Augmenting a Deep Learning Test Set to Increase its Mutation ScoreVincenzo Riccio, Nargiz Humbatova, Gunel Jahangirova, Paolo TonellaASE 2021 · 被引用 41 次
- ACETest: Automated Constraint Extraction for Testing Deep Learning OperatorsJingyi Shi, Yang Xiao, Yuekang Li, Yeting Li 等ISSTA 2023 · 被引用 24 次
- Repairing Failure-inducing Inputs with Input ReflectionYan Xiao, Yun Lin, Ivan Beschastnikh, Changsheng Sun 等ASE 2022 · 被引用 8 次
- Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative ModelsTong Che, Xiaofeng Liu, Site Li, Yubin Ge 等AAAI 2021 · 被引用 54 次
