On Focal Loss for Class-Posterior Probability Estimation: A Theoretical Perspective
Nontawat Charoenphakdee, Jayakorn Vongkulbhisal, Nuttapong Chairatanakul, Masashi Sugiyama
Abstract
The focal loss has demonstrated its effectiveness in many real-world applications such as object detection and image classification, but its theoretical understanding has been limited so far. In this paper, we first prove that the focal loss is classification-calibrated, i.e., its minimizer surely yields the Bayes-optimal classifier and thus the use of the focal loss in classification can be theoretically justified. However, we also prove a negative fact that the focal loss is not strictly proper, i.e., the confidence score of the classifier obtained by focal loss minimization does not match the true class-posterior probability and thus it is not reliable as a class-posterior probability estimator. To mitigate this problem, we next prove that a particular closed-form transformation of the confidence score allows us to recover the true class-posterior probability. Through experiments on benchmark datasets, we demonstrate that our proposed transformation significantly improves the accuracy of class-posterior probability estimation.
*Nontawat and Jayakorn contributed equally.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aef6fdbf-be5a-40f5-b087-c82d881be871Cited by top-tier papers6
- Dual Focal Loss for CalibrationLinwei Tao, Minjing Dong, Chang XuICML 2023 · 56 citations
- Generalizing Consistent Multi-Class Classification with Rejection to be Compatible with Arbitrary LossesYuzhou Cao, Tianchi Cai, Lei Feng, Lihong Gu et al.NeurIPS 2022 · 42 citations
- Learning to Find Good Models in RANSACDaniel Barath, Luca Cavalli, Marc PollefeysCVPR 2022 · 41 citations
- Towards Calibrated Multi-Label Deep Neural NetworksJiacheng Cheng, Nuno VasconcelosCVPR 2024
- Beyond One-Hot Labels: Semantic Mixing for Model CalibrationHaoyang Luo, Linwei Tao, Minjing Dong, Chang XuICML 2025
Builds on4
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- Peer Loss Functions: Learning from Noisy Labels without Knowing Noise RatesYang Liu, Hongyi GuoICML 2020 · 280 citations
- Consistent Estimators for Learning to Defer to an ExpertHussein Mozannar, David A. SontagICML 2020 · 267 citations
- Classification with Rejection Based on Cost-sensitive ClassificationNontawat Charoenphakdee, Zhenghang Cui, Yivan Zhang, Masashi SugiyamaICML 2021 · 78 citations
Related papers
- Uncertainty Weighted Gradients for Model CalibrationJinxu Lin, Linwei Tao, Minjing Dong, Chang XuCVPR 2025
- AdaFocal: Calibration-aware Adaptive Focal LossArindam Ghosh, Thomas Schaaf, Matthew GormleyNeurIPS 2022 · 71 citations
- Improved Balanced Classification with Theoretically Grounded Loss FunctionsCorinna Cortes, Mehryar Mohri, Yutao ZhongNeurIPS 2025 · 19 citations
- Towards Calibrated Model for Long-Tailed Visual Recognition from Prior PerspectiveZhengzhuo Xu, Zenghao Chai, Chun YuanNeurIPS 2021 · 77 citations
- Beyond calibration: estimating the grouping loss of modern neural networksAlexandre Perez-Lebel, Marine Le Morvan, Gaël VaroquauxICLR 2023 · 5 citations
