Influences on LLM Calibration: A Study of Response Agreement, Loss Functions, and Prompt Styles
Yuxi Xia, Pedro Henrique Luz de Araujo, Klim Zaporojets, Benjamin Roth
摘要
Calibration, the alignment between model confidence and prediction accuracy, is critical for the reliable deployment of large language models (LLMs). Existing works neglect to measure the generalization of their methods to other prompt styles and different sizes of LLMs. To address this, we define a controlled experimental setting covering 12 LLMs and four prompt styles. We additionally investigate if incorporating the response agreement of multiple LLMs and an appropriate loss function can improve calibration performance. Concretely, we build Calib-n, a novel framework that trains an auxiliary model for confidence estimation that aggregates responses from multiple LLMs to capture inter-model agreement. To optimize calibration, we integrate focal and AUC surrogate losses alongside binary cross-entropy. Experiments across four datasets demonstrate that both response agreement and focal loss improve calibration from baselines. We find that few-shot prompts are the most effective for auxiliary model-based methods, and auxiliary models demonstrate robust calibration performance across accuracy variations, outperforming LLMs' internal probabilities and verbalized confidences. These insights deepen the understanding of influence factors in LLM calibration, supporting their reliable deployment in diverse applications. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- BaseCal: Unsupervised Confidence Calibration via Base Model SignalsHexiang Tan, Wanli Yang, Junwei Zhang, Xin Chen 等ACL 2026 · 被引用 3 次
- EpiCaR: Knowing What You Don't Know Matters for Better Reasoning in LLMsJe Won Yeom, Jaewon Sok, Seonghyeon Park, Jeongjae Park 等ACL 2026 · 被引用 1 次
- Label and Explanation Variation in LLM-Based Annotation: a Case Study in Natural Language InferenceArtur Kulmizev, Erika Lombart, Patrick Watrin, Marie-Catherine de MarneffeACL 2026
它引用的顶会 Paper5
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
- Large-scale Robust Deep AUC Maximization: A New Surrogate Loss and Empirical Studies on Medical Image ClassificationZhuoning Yuan, Yan Yan, Milan Sonka, Tianbao YangICCV 2021 · 被引用 147 次
- Label Words are Anchors: An Information Flow Perspective for Understanding In-Context LearningLean Wang, Lei Li, Damai Dai, Deli Chen 等EMNLP 2023 · 被引用 26 次
- Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech DetectionMin Zhang, Jianfeng He, Taoran Ji, Chang-Tien LuACL 2024
相关 Paper
- Multicalibration for Confidence Scoring in LLMsGianluca Detommaso, Martin Bertran Lopez, Riccardo Fogliato, Aaron RothICML 2024 · 被引用 39 次
- Large Language Models Must Be Taught to Know What They Don't KnowSanyam Kapoor, Nate Gruver, Manley Roberts, Katie Collins 等NeurIPS 2024 · 被引用 124 次
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMsGabrielle Kaili-May Liu, Gal Yona, Avi Caciularu, Idan Szpektor 等EMNLP 2025
- Calibrating Large Language Models Using Their Generations OnlyDennis Ulmer, Martin Gubri, Hwaran Lee, Sangdoo Yun 等ACL 2024
- SteerConf: Steering LLMs for Confidence ElicitationZiang Zhou, Tianyuan Jin, Jieming Shi, Qing LiNeurIPS 2025 · 被引用 23 次
