When High Accuracy Hides Poor Calibration: Rethinking Confidence Evaluation in Transformer-Based Text Classification with Balanced Brier Score
Guilherme Fonseca, Gabriel Prenassi, Washington Cunha, Leonardo Chaves Dutra da Rocha, Marcos André Gonçalves
Abstract
Transformer-based Small (SLMs) and Large Language Models (LLMs) achieve strong effectiveness in text classification (TC), yet deployment requires reliable confidence estimates. Although miscalibration in Transformers has been reported, evidence for TC under fine-tuning remains limited. We evaluate the calibration of fine-tuned SLMs and LLMs against Logistic Regression, a classical, well-calibrated baseline, and find that, despite superior effectiveness, Transformers remain markedly overconfident. Crucially, we show that widely used calibration metrics, such as Expected Calibration Error and Brier Score, become biased in high-effectiveness regimes, where the dominance of correct predictions masks severe miscalibration on errors, sometimes even suggesting better calibration than Logistic Regression, a well-known calibrated method. To address this limitation, we propose the Balanced Brier Score (BBS), which balances the contribution of correct and incorrect predictions within confidence bins. BBS reveals substantially poorer calibration in both SLMs and LLMs, consistent with qualitative evidence from calibration curves. These findings challenge current calibration assessment practices and provide a more reliable alternative for evaluating confidence quality, particularly in high-effectiveness regimes where miscalibration may otherwise be underestimated.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Calibrating Large Language Models with Sample ConsistencyQing Lyu, Kumar Shridhar, Chaitanya Malaviya, Li Zhang et al.AAAI 2025 · 72 citations
- How Flawed Is ECE? An Analysis via Logit SmoothingMuthu Chidambaram, Holden Lee, Colin McSwiggen, Semon RezchikovICML 2024 · 7 citations
- Instance-Selection-Inspired Undersampling Strategies for Bias Reduction in Small and Large Language Models for Binary Text ClassificationGuilherme Fonseca, Washington Cunha, Gabriel Prenassi, Marcos André Gonçalves et al.ACL 2025
- Calibrating Large Language Models Using Their Generations OnlyDennis Ulmer, Martin Gubri, Hwaran Lee, Sangdoo Yun et al.ACL 2024
Related papers
- ConfTuner: Training Large Language Models to Express Their Confidence VerballyYibo Li, Miao Xiong, Jiaying Wu, Bryan HooiNeurIPS 2025 · 43 citations
- Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language ModelsDavid Bani-Harouni, Chantal Pellegrini, Paul Stangel, Ege Özsoy et al.ICLR 2026 · 49 citations
- Confidence Calibration of Classifiers with Many ClassesAdrien Le-Coz, Stéphane Herbin, Faouzi AdjedNeurIPS 2024 · 21 citations
- Calibrating Zero-shot Cross-lingual (Un-)structured PredictionsZhengping Jiang, Anqi Liu, Benjamin Van DurmeEMNLP 2022 · 4 citations
- Preserving Pre-trained Features Helps Calibrate Fine-tuned Language ModelsGuande He, Jianfei Chen, Jun ZhuICLR 2023 · 1 citation
