Calibration Across Layers: Understanding Calibration Evolution in LLMs
Abhinav Joshi, Areeb Ahmad, Ashutosh Modi
Abstract
Large Language Models (LLMs) have demonstrated inherent calibration capabilities, where predicted probabilities align well with correctness, despite prior findings that deep neural networks are often overconfident. Recent studies have linked this behavior to specific components in the final layer, such as entropy neurons and the unembedding matrix's null space. In this work, we provide a complementary perspective by investigating how calibration evolves throughout the network's depth. Analyzing multiple open-weight models on the MMLU benchmark, we uncover a distinct confidence correction phase in the upper/later layers, where model confidence is actively recalibrated after decision certainty has been reached. Furthermore, we identify a low-dimensional calibration direction in the residual stream whose perturbation significantly improves calibration metrics (ECE and MCE) without harming accuracy. Our findings suggest that calibration is a distributed phenomenon, shaped throughout the network's forward pass, not just in its final projection, providing new insights into how confidence-regulating mechanisms operate within LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5c282fb-7159-44e6-ba85-d4c5a735aa6bCited by top-tier papers4
- Geometry of Decision Making in Language ModelsAbhinav Joshi, Divyanshu Bhatt, Ashutosh ModiNeurIPS 2025 · 12 citations
- Beyond Components: Singular Vector-Based Interpretability of Transformer CircuitsAreeb Ahmad, Abhinav Joshi, Ashutosh ModiNeurIPS 2025 · 9 citations
- RSTR: Reducing SpatioTemporal Redundancy in Diffusion TransformersRuitong Sun, Tianze Yang, Wei Niu, Jin SunICML 2026 · 2 citations
- Correctness-Optimized Residual Activation Lens (CORAL): Transferrable and Calibration-Aware Inference-Time SteeringMiranda Miao, Young Min Cho, Lyle UngarICML 2026
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank ReductionPratyusha Sharma, Jordan T. Ash, Dipendra MisraICLR 2024 · 135 citations
- Calibrating Large Language Models with Sample ConsistencyQing Lyu, Kumar Shridhar, Chaitanya Malaviya, Li Zhang et al.AAAI 2025 · 72 citations
Related papers
- Calibrating LLM Confidence by Probing Perturbed Representation StabilityReza Khanmohammadi, Erfan Miahi, Mehrsa Mardikoraem, Simerjot Kaur et al.EMNLP 2025 · 1 citation
- BaseCal: Unsupervised Confidence Calibration via Base Model SignalsHexiang Tan, Wanli Yang, Junwei Zhang, Xin Chen et al.ACL 2026 · 3 citations
- Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMsPreetum Nakkiran, Arwen Bradley, Adam Golinski, Eugène Ndiaye et al.ICLR 2026 · 17 citations
- Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal BeliefZeguan Xiao, Diyang Dou, Boya Xiong, Yun Chen et al.AAAI 2026
- A Close Look into the Calibration of Pre-trained Language ModelsYangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu et al.ACL 2023 · 12 citations
