Lune

ICML2026顶会

Precision-Induced Miscalibration: Understanding and Correcting Confidence Distortion in Quantized Neural Networks

Jiawei Gu, Fengyuan Nie, Hao Tang, Yanpeng Sun

出版方
2026年份

摘要

Low-precision arithmetic is pervasive in neural network training and deployment, yet its effect on prediction confidence, not just accuracy, remains unexamined. We show that the softmax function amplifies logit-space quantization errors in an input-dependent manner: confidence distortion scales with the product of precision-dependent error bound ϵ and logit norm, peaking when the model is confident but not saturated. This explains why identical models report different confidence values across precisions, a phenomenon we term Precision Split. During training, the same mechanism causes gradient underflow: when logit margins exceed a precision-dependent threshold, gradients vanish and samples silently stop contributing to learning. Since logit norm serves as a computable proxy for precision-induced risk, we propose Precision-Aware Confidence Scaling (PACS), which applies sample-adaptive temperature inversely related to this risk, with sub-onepercent overhead and no full-precision computation required. On ImageNet with mixed-precision ResNet-50, PACS reduces Expected Calibration Error from 5.82% to 1.92% while maintaining accuracy, with consistent improvements across architectures, precision formats, and modalities. ✓ ✗ 1.2× Unit Scaling † 76.45 5.65 17.8 0.938 ✗ ✓ retrain PACS 76.13 1.92 7.6 0.871 ✗ ✗ 1.0004× PACS + LS † 76.52 1.54 6.2 0.852 ✗ ✓ 1.0004× PACS + LN 76.13 1.75 6.8 0.862 ✓ ✗ 1.2×

including occasional NaN losses (ablation in Section 4.3).

Parameter interpretation. The threshold τ controls when PACS intervenes; α controls how sharply it transitions; λ controls how much correction is applied. Values (τ = 0.01, α = 0.005, λ = 1.0) derive from numerical analysis rather than tuning: τ = ϵ × 10 triggers intervention when s > 10; λ = 1 yields T max = 2, sufficient to pull dangerous margins below saturation thresholds (Appendix B.10).

Prediction invariance. Since T (x) > 0 is a positive scalar, applying PACS to a fixed logit vector does not change the predicted class:

Therefore, when PACS is used as an inference-time postprocessing method on fixed logits, top-1 and top-5 accuracy remain exactly unchanged. If PACS is inserted during training, the final trained model may obtain different accuracy because the optimization trajectory changes

Computational overhead. PACS adds three elementary operations per sample: one reduction (abs().max()), one sigmoid, and one division. These incur <0.04% wall-clock overhead relative to the much more expensive trunk forward passes. Memory overhead is three scalars per sample. Method Training PPL↓ ECE↓ Cost Standard BF16 ✗ 8.34 12.3% 1.0× Soft-Prompt ✓ 8.25 7.6% 1.1×+train PACS ✗ 8.31 7.8% 1.004× PACS + Soft-Prompt ✓ 8.23 6.5% 1.1×+train Reliability diagrams. Figure 3 visualizes calibration quality. Standard FP16 exhibits systematic overconfidence (bars above diagonal). Logit Normalization (LN) improves calibration but retains residual bias in mid-confidence bins.

PACS achieves the closest alignment to perfect calibration across all confidence levels, with particularly strong correction in the critical 0.7-0.9 range identified in Section 2.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖