When High Accuracy Hides Poor Calibration: Rethinking Confidence Evaluation in Transformer-Based Text Classification with Balanced Brier Score
Guilherme Fonseca, Gabriel Prenassi, Washington Cunha, Leonardo Chaves Dutra da Rocha, Marcos André Gonçalves
摘要
Transformer-based Small (SLMs) and Large Language Models (LLMs) achieve strong effectiveness in text classification (TC), yet deployment requires reliable confidence estimates. Although miscalibration in Transformers has been reported, evidence for TC under fine-tuning remains limited. We evaluate the calibration of fine-tuned SLMs and LLMs against Logistic Regression, a classical, well-calibrated baseline, and find that, despite superior effectiveness, Transformers remain markedly overconfident. Crucially, we show that widely used calibration metrics, such as Expected Calibration Error and Brier Score, become biased in high-effectiveness regimes, where the dominance of correct predictions masks severe miscalibration on errors, sometimes even suggesting better calibration than Logistic Regression, a well-known calibrated method. To address this limitation, we propose the Balanced Brier Score (BBS), which balances the contribution of correct and incorrect predictions within confidence bins. BBS reveals substantially poorer calibration in both SLMs and LLMs, consistent with qualitative evidence from calibration curves. These findings challenge current calibration assessment practices and provide a more reliable alternative for evaluating confidence quality, particularly in high-effectiveness regimes where miscalibration may otherwise be underestimated.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Calibrating Large Language Models with Sample ConsistencyQing Lyu, Kumar Shridhar, Chaitanya Malaviya, Li Zhang 等AAAI 2025 · 被引用 72 次
- How Flawed Is ECE? An Analysis via Logit SmoothingMuthu Chidambaram, Holden Lee, Colin McSwiggen, Semon RezchikovICML 2024 · 被引用 7 次
- Instance-Selection-Inspired Undersampling Strategies for Bias Reduction in Small and Large Language Models for Binary Text ClassificationGuilherme Fonseca, Washington Cunha, Gabriel Prenassi, Marcos André Gonçalves 等ACL 2025
- Calibrating Large Language Models Using Their Generations OnlyDennis Ulmer, Martin Gubri, Hwaran Lee, Sangdoo Yun 等ACL 2024
相关 Paper
- ConfTuner: Training Large Language Models to Express Their Confidence VerballyYibo Li, Miao Xiong, Jiaying Wu, Bryan HooiNeurIPS 2025 · 被引用 43 次
- Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language ModelsDavid Bani-Harouni, Chantal Pellegrini, Paul Stangel, Ege Özsoy 等ICLR 2026 · 被引用 49 次
- Confidence Calibration of Classifiers with Many ClassesAdrien Le-Coz, Stéphane Herbin, Faouzi AdjedNeurIPS 2024 · 被引用 21 次
- Calibrating Zero-shot Cross-lingual (Un-)structured PredictionsZhengping Jiang, Anqi Liu, Benjamin Van DurmeEMNLP 2022 · 被引用 4 次
- Preserving Pre-trained Features Helps Calibrate Fine-tuned Language ModelsGuande He, Jianfei Chen, Jun ZhuICLR 2023 · 被引用 1 次
