Testing DNN image classifiers for confusion & bias errors
Yuchi Tian, Ziyuan Zhong, Vicente Ordonez, Gail E. Kaiser, Baishakhi Ray
Abstract
Image classifiers are an important component of today's software, from consumer and business applications to safety-critical domains. The advent of Deep Neural Networks (DNNs) is the key catalyst behind such wide-spread success. However, wide adoption comes with serious concerns about the robustness of software systems dependent on DNNs for image classification, as several severe erroneous behaviors have been reported under sensitive and critical circumstances. We argue that developers need to rigorously test their software's image classifiers and delay deployment until acceptable. We present an approach to testing image classifier robustness based on class property violations. We found that many of the reported erroneous cases in popular DNN image classifiers occur because the trained models confuse one class with another or show biases towards some classes over others. These bugs usually violate some class properties of one or more of those classes. Most DNN testing techniques focus on perimage violations, so fail to detect class-level confusions or biases. We developed a testing technique to automatically detect classbased confusion and bias errors in DNN-driven image classification software. We evaluated our implementation, DeepInspect, on several popular image classifiers with precision up to 100% (avg. 72.6%) for confusion errors, and up to 84.3% (avg. 66.8%) for bias errors. DeepInspect found hundreds of classification mistakes in widelyused models, many exposing errors indicating confusion or bias. CCS CONCEPTS • Software and its engineering → Software testing and debugging; • Computing methodologies → Neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b7ba57e-a930-4521-8f57-ae3a59ad5302Cited by top-tier papers5
- Are My Deep Learning Systems Fair? An Empirical Study of Fixed-Seed TrainingShangshu Qian, Hung Viet Pham, Thibaud Lutellier, Zeou Hu et al.NeurIPS 2021 · 49 citations
- ModelDiff: testing-based DNN similarity comparison for model reuse detectionYuanchun Li, Ziqi Zhang, Bingyan Liu, Ziyue Yang et al.ISSTA 2021 · 44 citations
- Automated testing of image captioning systemsBoxi Yu, Zhiqing Zhong, Xinran Qin, Jiayi Yao et al.ISSTA 2022 · 24 citations
- Coverage-Based Harmfulness Testing for LLM Code TransformationHonghao Tan, Haibo Wang, Diany Pressato, Yisen Xu et al.ASE 2025 · 2 citations
- FAST: Boosting Uncertainty-based Test Prioritization Methods for Neural Networks via Feature SelectionJialuo Chen, Jingyi Wang, Xiyue Zhang, Youcheng Sun et al.ASE 2024
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- Formal Security Analysis of Neural Networks using Symbolic IntervalsShiqi Wang, Kexin Pei, Justin Whitehouse, Junfeng Yang et al.USENIX Security 2018 · 523 citations
Related papers
- DeepGini: prioritizing massive tests to enhance the robustness of deep neural networksYang Feng, Qingkai Shi, Xinyu Gao, Jun Wan et al.ISSTA 2020 · 206 citations
- DeepSample: DNN sampling-based testing for operational accuracy assessmentAntonio Guerriero, Roberto Pietrantuono, Stefano RussoICSE 2024 · 6 citations
- Distribution-Aware Testing of Neural Networks Using Generative ModelsSwaroopa Dola, Matthew B. Dwyer, Mary Lou SoffaICSE 2021 · 3 citations
- DeepDiagnosis: Automatically Diagnosing Faults and Recommending Actionable Fixes in Deep Learning ProgramsMohammad Wardat, Breno Dantas Cruz, Wei Le, Hridesh RajanICSE 2022 · 46 citations
- AUTOTRAINER: An Automatic DNN Training Problem Detection and Repair SystemXiaoyu Zhang, Juan Zhai, Shiqing Ma, Chao ShenICSE 2021 · 62 citations
