Influences on LLM Calibration: A Study of Response Agreement, Loss Functions, and Prompt Styles
Yuxi Xia, Pedro Henrique Luz de Araujo, Klim Zaporojets, Benjamin Roth
Abstract
Calibration, the alignment between model confidence and prediction accuracy, is critical for the reliable deployment of large language models (LLMs). Existing works neglect to measure the generalization of their methods to other prompt styles and different sizes of LLMs. To address this, we define a controlled experimental setting covering 12 LLMs and four prompt styles. We additionally investigate if incorporating the response agreement of multiple LLMs and an appropriate loss function can improve calibration performance. Concretely, we build Calib-n, a novel framework that trains an auxiliary model for confidence estimation that aggregates responses from multiple LLMs to capture inter-model agreement. To optimize calibration, we integrate focal and AUC surrogate losses alongside binary cross-entropy. Experiments across four datasets demonstrate that both response agreement and focal loss improve calibration from baselines. We find that few-shot prompts are the most effective for auxiliary model-based methods, and auxiliary models demonstrate robust calibration performance across accuracy variations, outperforming LLMs' internal probabilities and verbalized confidences. These insights deepen the understanding of influence factors in LLM calibration, supporting their reliable deployment in diverse applications. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a7af871-5e92-4115-98bc-08dc7495a37bCited by top-tier papers3
- BaseCal: Unsupervised Confidence Calibration via Base Model SignalsHexiang Tan, Wanli Yang, Junwei Zhang, Xin Chen et al.ACL 2026 · 3 citations
- EpiCaR: Knowing What You Don't Know Matters for Better Reasoning in LLMsJe Won Yeom, Jaewon Sok, Seonghyeon Park, Jeongjae Park et al.ACL 2026 · 1 citation
- Label and Explanation Variation in LLM-Based Annotation: a Case Study in Natural Language InferenceArtur Kulmizev, Erika Lombart, Patrick Watrin, Marie-Catherine de MarneffeACL 2026
Builds on5
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Large-scale Robust Deep AUC Maximization: A New Surrogate Loss and Empirical Studies on Medical Image ClassificationZhuoning Yuan, Yan Yan, Milan Sonka, Tianbao YangICCV 2021 · 147 citations
- Label Words are Anchors: An Information Flow Perspective for Understanding In-Context LearningLean Wang, Lei Li, Damai Dai, Deli Chen et al.EMNLP 2023 · 26 citations
- Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech DetectionMin Zhang, Jianfeng He, Taoran Ji, Chang-Tien LuACL 2024
Related papers
- Multicalibration for Confidence Scoring in LLMsGianluca Detommaso, Martin Bertran Lopez, Riccardo Fogliato, Aaron RothICML 2024 · 39 citations
- Large Language Models Must Be Taught to Know What They Don't KnowSanyam Kapoor, Nate Gruver, Manley Roberts, Katie Collins et al.NeurIPS 2024 · 124 citations
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMsGabrielle Kaili-May Liu, Gal Yona, Avi Caciularu, Idan Szpektor et al.EMNLP 2025
- Calibrating Large Language Models Using Their Generations OnlyDennis Ulmer, Martin Gubri, Hwaran Lee, Sangdoo Yun et al.ACL 2024
- SteerConf: Steering LLMs for Confidence ElicitationZiang Zhou, Tianyuan Jin, Jieming Shi, Qing LiNeurIPS 2025 · 23 citations
