Cats Are Not Fish: Deep Learning Testing Calls for Out-Of-Distribution Awareness
David Berend, Xiaofei Xie, Lei Ma, Lingjun Zhou, Yang Liu, Chi Xu, Jianjun Zhao
摘要
As Deep Learning (DL) is continuously adopted in many industrial applications, its quality and reliability start to raise concerns. Similar to the traditional software development process, testing the DL software to uncover its defects at an early stage is an effective way to reduce risks after deployment. According to the fundamental assumption of deep learning, the DL software does not provide statistical guarantee and has limited capability in handling data that falls outside of its learned distribution, i.e., out-of-distribution (OOD) data. Although recent progress has been made in designing novel testing techniques for DL software, which can detect thousands of errors, the current state-of-the-art DL testing techniques usually do not take the distribution of generated test data into consideration. It is therefore hard to judge whether the "identified errors" are indeed meaningful errors to the DL application (i.e., due to quality issues of the model) or outliers that cannot be handled by the current model (i.e., due to the lack of training data). Tofill this gap, we take the first step and conduct a large scale empirical study, with a total of 451 experiment configurations, 42 deep neural networks (DNNs) and 1.2 million test data instances, to investigate and characterize the impact of OOD-awareness on DL testing. We further analyze the consequences when DL systems go into production by evaluating the effectiveness of adversarial retraining with distribution-aware errors. The results confirm that introducing data distribution awareness in both testing and enhancement phases outperforms distribution unaware retraining by up to 21.5%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Are Machine Learning Cloud APIs Used Correctly?Chengcheng Wan, Shicheng Liu, Henry Hoffmann, Michael Maire 等ICSE 2021 · 被引用 37 次
- Unsupervised Layer-Wise Score Aggregation for Textual OOD DetectionMaxime Darrin, Guillaume Staerman, Eduardo Dadalto Câmara Gomes, Jackie C. K. Cheung 等AAAI 2024 · 被引用 18 次
- DistXplore: Distribution-Guided Testing for Evaluating and Enhancing Deep Learning SystemsLongtian Wang, Xiaofei Xie, Xiaoning Du, Meng Tian 等FSE 2023 · 被引用 15 次
- Distribution Models for Falsification and Verification of DNNsFelipe Toledo, David Shriver, Sebastian G. Elbaum, Matthew B. DwyerASE 2021 · 被引用 4 次
- Empirical Insights of Test Selection Metrics under Multiple Testing Objectives and Distribution ShiftsJingyu Zhang, Fan Wang, Jacky Keung, Yihan Liao 等FSE 2026
它引用的顶会 Paper4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Input Complexity and Out-of-distribution Detection with Likelihood-based Generative ModelsJoan Serrà, David Álvarez, Vicenç Gómez, Olga Slizovskaia 等ICLR 2020 · 被引用 307 次
- Towards characterizing adversarial defects of deep learning software from the lens of uncertaintyXiyue Zhang, Xiaofei Xie, Lei Ma, Xiaoning Du 等ICSE 2020 · 被引用 69 次
- Amora: Black-box Adversarial Morphing AttackRun Wang, Felix Juefei-Xu, Qing Guo, Yihao Huang 等ACM MM 2020 · 被引用 40 次
相关 Paper
- A Statistical Framework for Efficient Out of Distribution Detection in Deep Neural NetworksMatan Haroush, Tzviel Frostig, Ruth Heller, Daniel SoudryICLR 2022 · 被引用 40 次
- Repairing Failure-inducing Inputs with Input ReflectionYan Xiao, Yun Lin, Ivan Beschastnikh, Changsheng Sun 等ASE 2022 · 被引用 8 次
- Navigating the Testing of Evolving Deep Learning Systems: An Exploratory Interview StudyHanmo You, Zan Wang, Bin Lin, Junjie ChenICSE 2025 · 被引用 1 次
- Meta OOD Learning For Continuously Adaptive OOD DetectionXinheng Wu, Jie Lu, Zhen Fang, Guangquan ZhangICCV 2023 · 被引用 15 次
- Distribution-Aware Testing of Neural Networks Using Generative ModelsSwaroopa Dola, Matthew B. Dwyer, Mary Lou SoffaICSE 2021 · 被引用 3 次
