Multicalibration for Confidence Scoring in LLMs
Gianluca Detommaso, Martin Bertran Lopez, Riccardo Fogliato, Aaron Roth
Abstract
This paper proposes the use of "multicalibration" to yield interpretable and reliable confidence scores for outputs generated by large language models (LLMs). Multicalibration asks for calibration not just marginally, but simultaneously across various intersecting groupings of the data. We show how to form groupings for prompt/completion pairs that are correlated with the probability of correctness via two techniques: clustering within an embedding space, and "selfannotation" -querying the LLM by asking it various yes-or-no questions about the prompt. We also develop novel variants of multicalibration algorithms that offer performance improvements by reducing their tendency to overfit. Through systematic benchmarking across various question answering datasets and LLMs, we show how our techniques can yield confidence scores that provide substantial improvements in fine-grained measures of both calibration and accuracy compared to existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9520e99-116d-44cd-b488-eab81ef0c837Cited by top-tier papers16
- Large language model validity via enhanced conformal prediction methodsJohn J. Cherian, Isaac Gibbs, Emmanuel J. CandèsNeurIPS 2024 · 120 citations
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme et al.ACL 2024 · 27 citations
- When is Multicalibration Post-Processing Necessary?Dutch Hansen, Siddartha Devic, Preetum Nakkiran, Vatsal SharanNeurIPS 2024 · 21 citations
- Self-Healing Machine Learning: A Framework for Autonomous Adaptation in Real-World EnvironmentsPaulius Rauba, Nabeel Seedat, Krzysztof Kacprzyk, Mihaela van der SchaarNeurIPS 2024 · 15 citations
- Conformal Information Pursuit for Interactively Guiding Large Language ModelsKwan Ho Ryan Chan, Yuyan Ge, Edgar Dobriban, Hamed Hassani et al.NeurIPS 2025 · 9 citations
Builds on10
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts et al.ACL 2023 · 319 citations
- A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text GenerationTianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao et al.ACL 2022 · 194 citations
- Conformal Language ModelingVictor Quach, Adam Fisch, Tal Schuster, Adam Yala et al.ICLR 2024 · 132 citations
- Practical Adversarial Multivalid Conformal PredictionOsbert Bastani, Varun Gupta, Christopher Jung, Georgy Noarov et al.NeurIPS 2022 · 82 citations
Related papers
- QA-Calibration of Language Model Confidence ScoresPutra Manggala, Atalanti-Anastasia Mastakouri, Elke Kirschbaum, Shiva Prasad Kasiviswanathan et al.ICLR 2025
- Calibrating the Confidence of Large Language Models by Eliciting FidelityMozhi Zhang, Mianqiu Huang, Rundong Shi, Linsen Guo et al.EMNLP 2024 · 5 citations
- Calibrating Large Language Models Using Their Generations OnlyDennis Ulmer, Martin Gubri, Hwaran Lee, Sangdoo Yun et al.ACL 2024
- Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language ModelsDavid Bani-Harouni, Chantal Pellegrini, Paul Stangel, Ege Özsoy et al.ICLR 2026 · 49 citations
- Large Language Models Must Be Taught to Know What They Don't KnowSanyam Kapoor, Nate Gruver, Manley Roberts, Katie Collins et al.NeurIPS 2024 · 124 citations
