On Calibration of Pre-trained Code Models
Zhenhao Zhou, Chaofeng Sha, Xin Peng
摘要
Pre-trained code models have achieved notable success in the field of Software Engineering (SE). However, existing studies have predominantly focused on improving model performance, with limited attention given to other critical aspects such as model calibration. Model calibration, which refers to the accurate estimation of predictive uncertainty, is a vital consideration in practical applications. Therefore, in order to advance the understanding of model calibration in SE, we conduct a comprehensive investigation into the calibration of pre-trained code models in this paper. Our investigation focuses on five pre-trained code models and four code understanding tasks, including analyses of calibration in both in-distribution and out-of-distribution settings. Several key insights are uncovered: (1) pre-trained code models may suffer from the issue of over-confidence; (2) temperature scaling and label smoothing are effective in calibrating code models in in-distribution data; (3) the issue of over-confidence in pre-trained code models worsens in different out-of-distribution settings, and the effectiveness of temperature scaling and label smoothing diminishes. All materials used in our experiments are available at https://github.com/queserasera22/Calibration-of-Pretrained-Code-Models.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Calibration and Correctness of Language Models for CodeClaudio Spiess, David Gros, Kunal Suresh Pai, Michael Pradel 等ICSE 2025 · 被引用 21 次
- Decoding Secret Memorization in Code LLMs Through Token-Level CharacterizationYuqing Nie, Chong Wang, Kailong Wang, Guoai Xu 等ICSE 2025 · 被引用 10 次
- Calibration of Large Language Models on Code SummarizationYuvraj Virk, Premkumar T. Devanbu, Toufique AhmedFSE 2025 · 被引用 5 次
- Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language ModelsMoule Lin, Shuhao Guan, Andrea Patane, David Gregg 等ICML 2026
相关 Paper
- A Close Look into the Calibration of Pre-trained Language ModelsYangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu 等ACL 2023 · 被引用 12 次
- An Empirical Comparison of Pre-Trained Models of Source CodeChangan Niu, Chuanyi Li, Vincent Ng, Dongxiao Chen 等ICSE 2023 · 被引用 71 次
- Calibrating Zero-shot Cross-lingual (Un-)structured PredictionsZhengping Jiang, Anqi Liu, Benjamin Van DurmeEMNLP 2022 · 被引用 4 次
- Proximity-Informed Calibration for Deep Neural NetworksMiao Xiong, Ailin Deng, Pang Wei Koh, Jiaying Wu 等NeurIPS 2023 · 被引用 30 次
- Robust Calibration with Multi-domain Temperature ScalingYaodong Yu, Stephen Bates, Yi Ma, Michael I. JordanNeurIPS 2022 · 被引用 58 次
