When Does Optimizing a Proper Loss Yield Calibration?
Jaroslaw Blasiok, Parikshit Gopalan, Lunjia Hu, Preetum Nakkiran
摘要
Optimizing proper loss functions is popularly believed to yield predictors with good calibration properties; the intuition being that for such losses, the global optimum is to predict the ground-truth probabilities, which is indeed calibrated. However, typical machine learning models are trained to approximately minimize loss over restricted families of predictors, that are unlikely to contain the ground truth. Under what circumstances does optimizing proper loss over a restricted family yield calibrated models? What precise calibration guarantees does it give? In this work, we provide a rigorous answer to these questions. We replace the global optimality with a local optimality condition stipulating that the (proper) loss of the predictor cannot be reduced much by post-processing its predictions with a certain family of Lipschitz functions. We show that any predictor with this local optimality satisfies smooth calibration as defined in Kakade and Foster (2008); Błasiok et al. (2023b). Local optimality is plausibly satisfied by well-trained DNNs, which suggests an explanation for why they are calibrated from proper loss minimization alone. Finally, we show that the connection between local optimality and calibration error goes both ways: nearly calibrated predictors are also nearly locally optimal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Smooth ECE: Principled Reliability Diagrams via Kernel SmoothingJaroslaw Blasiok, Preetum NakkiranICLR 2024 · 被引用 59 次
- ConfTuner: Training Large Language Models to Express Their Confidence VerballyYibo Li, Miao Xiong, Jiaying Wu, Bryan HooiNeurIPS 2025 · 被引用 43 次
- When is Multicalibration Post-Processing Necessary?Dutch Hansen, Siddartha Devic, Preetum Nakkiran, Vatsal SharanNeurIPS 2024 · 被引用 21 次
- Experts Don't Cheat: Learning What You Don't Know By Predicting PairsDaniel D. Johnson, Daniel Tarlow, David Duvenaud, Chris J. MaddisonICML 2024 · 被引用 18 次
- Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMsPreetum Nakkiran, Arwen Bradley, Adam Golinski, Eugène Ndiaye 等ICLR 2026 · 被引用 17 次
它引用的顶会 Paper9
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz 等NeurIPS 2020 · 被引用 674 次
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis 等NeurIPS 2021 · 被引用 633 次
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
- Soft Calibration Objectives for Neural NetworksArchit Karandikar, Nicholas Cain, Dustin Tran, Balaji Lakshminarayanan 等NeurIPS 2021 · 被引用 127 次
- Smooth ECE: Principled Reliability Diagrams via Kernel SmoothingJaroslaw Blasiok, Preetum NakkiranICLR 2024 · 被引用 59 次
相关 Paper
- Smooth Calibration Error: Uniform Convergence and Functional Gradient AnalysisFutoshi Futami, Atsushi NitandaICLR 2026
- Omnipredictors for Constrained OptimizationLunjia Hu, Inbal Rachel Livni Navon, Omer Reingold, Chutong YangICML 2023 · 被引用 17 次
- Calibrated and Sharp Uncertainties in Deep Learning via Density EstimationVolodymyr Kuleshov, Shachi DeshpandeICML 2022 · 被引用 44 次
- Efficient Calibration for Decision MakingParikshit Gopalan, Konstantinos Stavropoulos, Kunal Talwar, Pranay TankalaSTOC 2026 · 被引用 3 次
- From Individual Calibration to Reliable Classifiers: ALD Parameterization with mPAIC GuaranteesDeming Sheng, Ricardo HenaoICML 2026
