Cats Are Not Fish: Deep Learning Testing Calls for Out-Of-Distribution Awareness
David Berend, Xiaofei Xie, Lei Ma, Lingjun Zhou, Yang Liu, Chi Xu, Jianjun Zhao
Abstract
As Deep Learning (DL) is continuously adopted in many industrial applications, its quality and reliability start to raise concerns. Similar to the traditional software development process, testing the DL software to uncover its defects at an early stage is an effective way to reduce risks after deployment. According to the fundamental assumption of deep learning, the DL software does not provide statistical guarantee and has limited capability in handling data that falls outside of its learned distribution, i.e., out-of-distribution (OOD) data. Although recent progress has been made in designing novel testing techniques for DL software, which can detect thousands of errors, the current state-of-the-art DL testing techniques usually do not take the distribution of generated test data into consideration. It is therefore hard to judge whether the "identified errors" are indeed meaningful errors to the DL application (i.e., due to quality issues of the model) or outliers that cannot be handled by the current model (i.e., due to the lack of training data). Tofill this gap, we take the first step and conduct a large scale empirical study, with a total of 451 experiment configurations, 42 deep neural networks (DNNs) and 1.2 million test data instances, to investigate and characterize the impact of OOD-awareness on DL testing. We further analyze the consequences when DL systems go into production by evaluating the effectiveness of adversarial retraining with distribution-aware errors. The results confirm that introducing data distribution awareness in both testing and enhancement phases outperforms distribution unaware retraining by up to 21.5%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Are Machine Learning Cloud APIs Used Correctly?Chengcheng Wan, Shicheng Liu, Henry Hoffmann, Michael Maire et al.ICSE 2021 · 37 citations
- Unsupervised Layer-Wise Score Aggregation for Textual OOD DetectionMaxime Darrin, Guillaume Staerman, Eduardo Dadalto Câmara Gomes, Jackie C. K. Cheung et al.AAAI 2024 · 18 citations
- DistXplore: Distribution-Guided Testing for Evaluating and Enhancing Deep Learning SystemsLongtian Wang, Xiaofei Xie, Xiaoning Du, Meng Tian et al.FSE 2023 · 15 citations
- Distribution Models for Falsification and Verification of DNNsFelipe Toledo, David Shriver, Sebastian G. Elbaum, Matthew B. DwyerASE 2021 · 4 citations
- Empirical Insights of Test Selection Metrics under Multiple Testing Objectives and Distribution ShiftsJingyu Zhang, Fan Wang, Jacky Keung, Yihan Liao et al.FSE 2026
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Input Complexity and Out-of-distribution Detection with Likelihood-based Generative ModelsJoan Serrà, David Álvarez, Vicenç Gómez, Olga Slizovskaia et al.ICLR 2020 · 307 citations
- Towards characterizing adversarial defects of deep learning software from the lens of uncertaintyXiyue Zhang, Xiaofei Xie, Lei Ma, Xiaoning Du et al.ICSE 2020 · 69 citations
- Amora: Black-box Adversarial Morphing AttackRun Wang, Felix Juefei-Xu, Qing Guo, Yihao Huang et al.ACM MM 2020 · 40 citations
Related papers
- A Statistical Framework for Efficient Out of Distribution Detection in Deep Neural NetworksMatan Haroush, Tzviel Frostig, Ruth Heller, Daniel SoudryICLR 2022 · 40 citations
- Repairing Failure-inducing Inputs with Input ReflectionYan Xiao, Yun Lin, Ivan Beschastnikh, Changsheng Sun et al.ASE 2022 · 8 citations
- Navigating the Testing of Evolving Deep Learning Systems: An Exploratory Interview StudyHanmo You, Zan Wang, Bin Lin, Junjie ChenICSE 2025 · 1 citation
- Meta OOD Learning For Continuously Adaptive OOD DetectionXinheng Wu, Jie Lu, Zhen Fang, Guangquan ZhangICCV 2023 · 15 citations
- Distribution-Aware Testing of Neural Networks Using Generative ModelsSwaroopa Dola, Matthew B. Dwyer, Mary Lou SoffaICSE 2021 · 3 citations
