Operational calibration: debugging confidence errors for DNNs in the field
Zenan Li, Xiaoxing Ma, Chang Xu, Jingwei Xu, Chun Cao, Jian Lu
Abstract
Trained DNN models are increasingly adopted as integral parts of software systems, but they often perform deficiently in the field. A particularly damaging problem is that DNN models often give false predictions with high confidence, due to the unavoidable slight divergences between operation data and training data. To minimize the loss caused by inaccurate confidence, operational calibration, i.e., calibrating the confidence function of a DNN classifier against its operation domain, becomes a necessary debugging step in the engineering of the whole system.
Operational calibration is difficult considering the limited budget of labeling operation data and the weak interpretability of DNN models. We propose a Bayesian approach to operational calibration that gradually corrects the confidence given by the model under calibration with a small number of labeled operation data deliberately selected from a larger set of unlabeled operation data. The approach is made effective and efficient by leveraging the locality of the learned representation of the DNN model and modeling the calibration as Gaussian Process Regression. Comprehensive experiments with various practical datasets and DNN models show that it significantly outperformed alternative methods, and in some difficult tasks it eliminated about 71% to 97% high-confidence (>0.9) errors with only about 10% of the minimal amount of labeled operation data needed for practical learning techniques to barely work.
• Software and its engineering → Software testing and debugging; • Computing methodologies → Neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73d9c814-1b90-4634-8aaf-18b5cda14903Cited by top-tier papers7
- Prioritizing Test Inputs for Deep Neural Networks via Mutation AnalysisZan Wang, Hanmo You, Junjie Chen, Yingyi Zhang et al.ICSE 2021 · 117 citations
- Are Machine Learning Cloud APIs Used Correctly?Chengcheng Wan, Shicheng Liu, Henry Hoffmann, Michael Maire et al.ICSE 2021 · 37 citations
- Calibration and Correctness of Language Models for CodeClaudio Spiess, David Gros, Kunal Suresh Pai, Michael Pradel et al.ICSE 2025 · 21 citations
- Lightweight Approaches to DNN Regression Error Reduction: An Uncertainty Alignment PerspectiveZenan Li, Maorun Zhang, Jingwei Xu, Yuan Yao et al.ICSE 2023 · 3 citations
- Using a Sledgehammer to Crack a Nut? Revisiting Automated Compiler Fault IsolationYibiao Yang, Qingyang Li, Maolin Sun, Jiangchang Wu et al.ICSE 2026 · 1 citation
Builds on1
Related papers
- Operation is the hardest teacher: estimating DNN accuracy looking for mispredictionsAntonio Guerriero, Roberto Pietrantuono, Stefano RussoICSE 2021 · 3 citations
- Taking a Step Back with KCal: Multi-Class Kernel-Based Calibration for Deep Neural NetworksZhen Lin, Shubhendu Trivedi, Jimeng SunICLR 2023 · 2 citations
- DeepSample: DNN sampling-based testing for operational accuracy assessmentAntonio Guerriero, Roberto Pietrantuono, Stefano RussoICSE 2024 · 6 citations
- Rethinking Calibration of Deep Neural Networks: Do Not Be Afraid of OverconfidenceDeng-Bao Wang, Lei Feng, Min-Ling ZhangNeurIPS 2021 · 177 citations
- Exploring the Uncertainty Properties of Neural Networks' Implicit Priors in the Infinite-Width LimitBen Adlam, Jaehoon Lee, Lechao Xiao, Jeffrey Pennington et al.ICLR 2021 · 3 citations
