Calibration and Correctness of Language Models for Code
Claudio Spiess, David Gros, Kunal Suresh Pai, Michael Pradel, Md. Rafiqul Islam Rabin, Amin Alipour, Susmit Jha, Prem Devanbu, Toufique Ahmed
Abstract
Machine learning models are widely used, but can also often be wrong. Users would benefit from a reliable indication of whether a given output from a given model should be trusted, so a rational decision can be made whether to use the output or not. For example, outputs can be associated with a confidence measure; if this confidence measure is strongly associated with likelihood of correctness, then the model is said to be well-calibrated. A well-calibrated confidence measure can serve as a basis for rational, graduated decision-making on how much review and care is needed when using generated code. Calibration has so far been studied in mostly non-generative (e.g., classification) settings, especially in software engineering. However, generated code can quite often be wrong: Given generated code, developers must decide whether to use directly, use after varying intensity of careful review, or discard model-generated code. Thus, calibration is vital in generative settings. We make several contributions. We develop a framework for evaluating the calibration of code-generating models. We consider several tasks, correctness criteria, datasets, and approaches, and find that, by and large, generative code models we test are not well-calibrated out of the box. We then show how calibration can be improved using standard methods, such as Platt scaling. Since Platt scaling relies on the prior availability of correctness data, we evaluate the applicability and generalizability of Platt scaling in software engineering, discuss settings where it has good potential for practical use, and settings where it does not. Our contributions will lead to better-calibrated decision-making in the current use of code generated by language models, and offers a framework for future research to further improve calibration methods for generative models in software engineering.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7be7b202-0caa-4511-a5f6-a12fdaed0970Cited by top-tier papers12
- LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and MitigationZiyao Zhang, Chong Wang, Yanlin Wang, Ensheng Shi et al.ISSTA 2025 · 53 citations
- SQLens: An End-to-End Framework for Error Detection and Correction in Text-to-SQLYue Gong, Chuan Lei, Xiao Qin, Kapil Vaidya et al.NeurIPS 2025 · 21 citations
- Do Large Language Models Know What They Are Capable Of?Casey O. Barkan, Sidney Black, Oliver SourbutICLR 2026 · 11 citations
- Conformal Information Pursuit for Interactively Guiding Large Language ModelsKwan Ho Ryan Chan, Yuyan Ge, Edgar Dobriban, Hamed Hassani et al.NeurIPS 2025 · 9 citations
- Uncertainty-aware Generative RecommendationChenxiao Fan, Chongming Gao, Yaxin Gong, Haoyan Liu et al.KDD 2026 · 2 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis et al.NeurIPS 2021 · 633 citations
Related papers
- Calibrating LLMs for Text-to-SQL Parsing by Leveraging Sub-clause FrequenciesTerrance Liu, Shuyi Wang, Daniel Preotiuc-Pietro, Yash Chandarana et al.EMNLP 2025
- On Calibration of Pre-trained Code ModelsZhenhao Zhou, Chaofeng Sha, Xin PengICSE 2024 · 3 citations
- Probability Calibration for Knowledge Graph Embedding ModelsPedro Tabacof, Luca CostabelloICLR 2020 · 49 citations
- QA-Calibration of Language Model Confidence ScoresPutra Manggala, Atalanti-Anastasia Mastakouri, Elke Kirschbaum, Shiva Prasad Kasiviswanathan et al.ICLR 2025
- Uncertainty in Language Models: Assessment through Rank-CalibrationXinmeng Huang, Shuo Li, Mengxin Yu, Matteo Sesia et al.EMNLP 2024 · 9 citations
