Precision-Induced Miscalibration: Understanding and Correcting Confidence Distortion in Quantized Neural Networks
Jiawei Gu, Fengyuan Nie, Hao Tang, Yanpeng Sun
Abstract
Low-precision arithmetic is pervasive in neural network training and deployment, yet its effect on prediction confidence, not just accuracy, remains unexamined. We show that the softmax function amplifies logit-space quantization errors in an input-dependent manner: confidence distortion scales with the product of precision-dependent error bound ϵ and logit norm, peaking when the model is confident but not saturated. This explains why identical models report different confidence values across precisions, a phenomenon we term Precision Split. During training, the same mechanism causes gradient underflow: when logit margins exceed a precision-dependent threshold, gradients vanish and samples silently stop contributing to learning. Since logit norm serves as a computable proxy for precision-induced risk, we propose Precision-Aware Confidence Scaling (PACS), which applies sample-adaptive temperature inversely related to this risk, with sub-onepercent overhead and no full-precision computation required. On ImageNet with mixed-precision ResNet-50, PACS reduces Expected Calibration Error from 5.82% to 1.92% while maintaining accuracy, with consistent improvements across architectures, precision formats, and modalities. ✓ ✗ 1.2× Unit Scaling † 76.45 5.65 17.8 0.938 ✗ ✓ retrain PACS 76.13 1.92 7.6 0.871 ✗ ✗ 1.0004× PACS + LS † 76.52 1.54 6.2 0.852 ✗ ✓ 1.0004× PACS + LN 76.13 1.75 6.8 0.862 ✓ ✗ 1.2×
including occasional NaN losses (ablation in Section 4.3).
Parameter interpretation. The threshold τ controls when PACS intervenes; α controls how sharply it transitions; λ controls how much correction is applied. Values (τ = 0.01, α = 0.005, λ = 1.0) derive from numerical analysis rather than tuning: τ = ϵ × 10 triggers intervention when s > 10; λ = 1 yields T max = 2, sufficient to pull dangerous margins below saturation thresholds (Appendix B.10).
Prediction invariance. Since T (x) > 0 is a positive scalar, applying PACS to a fixed logit vector does not change the predicted class:
Therefore, when PACS is used as an inference-time postprocessing method on fixed logits, top-1 and top-5 accuracy remain exactly unchanged. If PACS is inserted during training, the final trained model may obtain different accuracy because the optimization trajectory changes
Computational overhead. PACS adds three elementary operations per sample: one reduction (abs().max()), one sigmoid, and one division. These incur <0.04% wall-clock overhead relative to the much more expensive trunk forward passes. Memory overhead is three scalars per sample. Method Training PPL↓ ECE↓ Cost Standard BF16 ✗ 8.34 12.3% 1.0× Soft-Prompt ✓ 8.25 7.6% 1.1×+train PACS ✗ 8.31 7.8% 1.004× PACS + Soft-Prompt ✓ 8.23 6.5% 1.1×+train Reliability diagrams. Figure 3 visualizes calibration quality. Standard FP16 exhibits systematic overconfidence (bars above diagonal). Logit Normalization (LN) improves calibration but retains residual bias in mid-confidence bins.
PACS achieves the closest alignment to perfect calibration across all confidence levels, with particularly strong correction in the critical 0.7-0.9 range identified in Section 2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e48e036-a4e1-4681-9835-86effae85bcbBuilds on12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis et al.NeurIPS 2021 · 633 citations
Related papers
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- Sample Margin-Aware Recalibration of Temperature ScalingHaolan Guo, Linwei Tao, Haoyang Luo, Minjing Dong et al.ICML 2026 · 3 citations
- Scaling Laws for PrecisionTanishq Kumar, Zachary Ankner, Benjamin Frederick Spector, Blake Bordelon et al.ICLR 2025
- Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMsBinxing Xu, Hao Gu, Lujun Li, Hao Wang et al.ACL 2026 · 2 citations
- Any-Precision Deep Neural NetworksHaichao Yu, Haoxiang Li, Humphrey Shi, Thomas S. Huang et al.AAAI 2021 · 79 citations
