Calibrating LLM Confidence by Probing Perturbed Representation Stability
Reza Khanmohammadi, Erfan Miahi, Mehrsa Mardikoraem, Simerjot Kaur, Ivan Brugere, Charese Smiley, Kundan Thind, Mohammad M. Ghassemi
Abstract
Miscalibration in Large Language Models (LLMs) undermines their reliability, highlighting the need for accurate confidence estimation. We introduce CCPS (Calibrating LLM Confidence by Probing Perturbed Representation Stability), a novel method analyzing internal representational stability in LLMs. CCPS applies targeted adversarial perturbations to final hidden states, extracts features reflecting the model's response to these perturbations, and uses a lightweight classifier to predict answer correctness. CCPS was evaluated on LLMs from 8B to 32B parameters (covering Llama, Qwen, and Mistral architectures) using MMLU and MMLU-Pro benchmarks in both multiplechoice and open-ended formats. Our results show that CCPS significantly outperforms current approaches. Across four LLMs and three MMLU variants, CCPS reduces Expected Calibration Error by approximately 55% and Brier score by 21%, while increasing accuracy by 5 percentage points, Area Under the Precision-Recall Curve by 4 percentage points, and Area Under the Receiver Operating Characteristic Curve by 6 percentage points, all relative to the strongest prior method. CCPS delivers an efficient, broadly applicable, and more accurate solution for estimating LLM confidence, thereby improving their trustworthiness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc ReasoningQi Cao, Shuhao Zhang, Ruizhe Zhou, Ruiyi Zhang et al.ICML 2026 · 2 citations
- Resisting Manipulative Bots in Meme Coin Copy Trading: A Multi-Agent Approach with Chain-of-Thought ReasoningYichen Luo, Yebo Feng, Jiahua Xu, Yang LiuWWW 2026 · 1 citation
- Margin-Adaptive Confidence Ranking for Reliable LLM JudgementGaojie Jin, Yong Tao, Lijia Yu, Tianjin HuangICML 2026
- CaliDist: Calibrating Large Language Models via Behavioral Robustness to DistractionMohammad Anas Jawad, Cornelia CarageaICML 2026
- Correctness-Optimized Residual Activation Lens (CORAL): Transferrable and Calibration-Aware Inference-Time SteeringMiranda Miao, Young Min Cho, Lyle UngarICML 2026
Builds on4
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- HaDeMiF: Hallucination Detection and Mitigation in Large Language ModelsXiaoling Zhou, Mingjie Zhang, Zhemg Lee, Wei Ye et al.ICLR 2025
Related papers
- Factual Confidence of LLMs: on Reliability and Robustness of Current EstimatorsMatéo Mahaut, Laura Aina, Paula Czarnowska, Momchil Hardalov et al.ACL 2024 · 7 citations
- Confidence Elicitation: A New Attack Vector for Large Language ModelsBrian Formento, Chuan-Sheng Foo, See-Kiong NgICLR 2025
- Learning to Route LLMs with Confidence TokensYu-Neng Chuang, Prathusha Kameswara Sarma, Parikshit Gopalan, John Boccio et al.ICML 2025
- Calibrating Large Language Models Using Their Generations OnlyDennis Ulmer, Martin Gubri, Hwaran Lee, Sangdoo Yun et al.ACL 2024
- Calibration Across Layers: Understanding Calibration Evolution in LLMsAbhinav Joshi, Areeb Ahmad, Ashutosh ModiEMNLP 2025 · 1 citation
