Testing DNN image classifiers for confusion & bias errors
Yuchi Tian, Ziyuan Zhong, Vicente Ordonez, Gail E. Kaiser, Baishakhi Ray
摘要
Image classifiers are an important component of today's software, from consumer and business applications to safety-critical domains. The advent of Deep Neural Networks (DNNs) is the key catalyst behind such wide-spread success. However, wide adoption comes with serious concerns about the robustness of software systems dependent on DNNs for image classification, as several severe erroneous behaviors have been reported under sensitive and critical circumstances. We argue that developers need to rigorously test their software's image classifiers and delay deployment until acceptable. We present an approach to testing image classifier robustness based on class property violations. We found that many of the reported erroneous cases in popular DNN image classifiers occur because the trained models confuse one class with another or show biases towards some classes over others. These bugs usually violate some class properties of one or more of those classes. Most DNN testing techniques focus on perimage violations, so fail to detect class-level confusions or biases. We developed a testing technique to automatically detect classbased confusion and bias errors in DNN-driven image classification software. We evaluated our implementation, DeepInspect, on several popular image classifiers with precision up to 100% (avg. 72.6%) for confusion errors, and up to 84.3% (avg. 66.8%) for bias errors. DeepInspect found hundreds of classification mistakes in widelyused models, many exposing errors indicating confusion or bias. CCS CONCEPTS • Software and its engineering → Software testing and debugging; • Computing methodologies → Neural networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Are My Deep Learning Systems Fair? An Empirical Study of Fixed-Seed TrainingShangshu Qian, Hung Viet Pham, Thibaud Lutellier, Zeou Hu 等NeurIPS 2021 · 被引用 49 次
- ModelDiff: testing-based DNN similarity comparison for model reuse detectionYuanchun Li, Ziqi Zhang, Bingyan Liu, Ziyue Yang 等ISSTA 2021 · 被引用 44 次
- Automated testing of image captioning systemsBoxi Yu, Zhiqing Zhong, Xinran Qin, Jiayi Yao 等ISSTA 2022 · 被引用 24 次
- Coverage-Based Harmfulness Testing for LLM Code TransformationHonghao Tan, Haibo Wang, Diany Pressato, Yisen Xu 等ASE 2025 · 被引用 2 次
- FAST: Boosting Uncertainty-based Test Prioritization Methods for Neural Networks via Feature SelectionJialuo Chen, Jingyi Wang, Xiyue Zhang, Youcheng Sun 等ASE 2024
它引用的顶会 Paper4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- Formal Security Analysis of Neural Networks using Symbolic IntervalsShiqi Wang, Kexin Pei, Justin Whitehouse, Junfeng Yang 等USENIX Security 2018 · 被引用 523 次
相关 Paper
- DeepGini: prioritizing massive tests to enhance the robustness of deep neural networksYang Feng, Qingkai Shi, Xinyu Gao, Jun Wan 等ISSTA 2020 · 被引用 206 次
- DeepSample: DNN sampling-based testing for operational accuracy assessmentAntonio Guerriero, Roberto Pietrantuono, Stefano RussoICSE 2024 · 被引用 6 次
- Distribution-Aware Testing of Neural Networks Using Generative ModelsSwaroopa Dola, Matthew B. Dwyer, Mary Lou SoffaICSE 2021 · 被引用 3 次
- DeepDiagnosis: Automatically Diagnosing Faults and Recommending Actionable Fixes in Deep Learning ProgramsMohammad Wardat, Breno Dantas Cruz, Wei Le, Hridesh RajanICSE 2022 · 被引用 46 次
- AUTOTRAINER: An Automatic DNN Training Problem Detection and Repair SystemXiaoyu Zhang, Juan Zhai, Shiqing Ma, Chao ShenICSE 2021 · 被引用 62 次
