Calibration of Large Language Models on Code Summarization
Yuvraj Virk, Premkumar T. Devanbu, Toufique Ahmed
摘要
A brief, fluent, and relevant summary can be helpful during program comprehension; however, such a summary does require significant human effort to produce. Often, good summaries are unavailable in software projects, which makes maintenance more difficult. There has been a considerable body of research into automated AI-based methods, using Large Language models (LLMs), to generate summaries of code; there also has been quite a bit of work on ways to measure the performance of such summarization methods, with special attention paid to how closely these AI-generated summaries resemble a summary a human might have produced. Measures such as BERTScore and BLEU have been suggested and evaluated with human-subject studies. However, LLM-generated summaries can be inaccurate, incomplete, etc: generally, too dissimilar to one that a good developer might write. Given an LLM-generated code summary, how can a user rationally judge if a summary is sufficiently good and reliable? Given just some input source code, and an LLM-generated summary, existing approaches can help judge brevity, fluency and relevance of the summary; however, it’s difficult to gauge whether an LLM-generated summary sufficiently resembles what a human might produce, without a “golden” human-produced summary to compare against. Prior research indicates that human-produced summaries are generally preferred by human-raters, so we explore this issue in this paper. We study this resemblance question as a calibration problem: given just the code & the summary from an LLM, can we compute a confidence measure, that provides a reliable indication of whether the summary sufficiently resembles what a human would have produced in this situation? We examine this question using several LLMs, for several languages, and in several different settings. Our investigation suggests approaches to provide reliable predictions of the likelihood that an LLM-generated summary would sufficiently resemble a summary a human might write for the same code.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language ModelsJessica Y. Bo, Sophia Wan, Ashton AndersonCHI 2025 · 被引用 31 次
- Retrofit: Continual Learning with Controlled Forgetting for Binary Security Detection and AnalysisYiling He, Junchi Lei, Hongyu She, Shuo Shao 等USENIX Security 2026
- Code-MUE: Measuring Code LLMs’ Uncertainty through Execution-Based Semantic Interaction GraphsXiaoning Ren, Yinxing Xue, Lei Ma, Yuheng HuangISSTA 2026
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 被引用 439 次
- Retrieval-Based Prompt Selection for Code-Related Few-Shot LearningNoor Nashid, Mifta Sintaha, Ali MesbahICSE 2023 · 被引用 156 次
相关 Paper
- SimLLM: Calculating Semantic Similarity in Code Summaries using a Large Language Model-Based ApproachXin Jin, Zhiqiang LinFSE 2024 · 被引用 8 次
- On Behavioral Alignment of Model-Code and Human-Code Understandability via Behavioral ProxiesXiaokai Rong, Aashish Yadavally, Hridya Dhulipala, Anh H. N. Nguyen 等ISSTA 2026
- Source Code Summarization in the Era of Large Language ModelsWeisong Sun, Yun Miao, Yuekang Li, Hongyu Zhang 等ICSE 2025 · 被引用 37 次
- Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?Bonan Kou, Shengmai Chen, Zhijie Wang, Lei Ma 等FSE 2024 · 被引用 8 次
- ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code SummarizationSuyoung Bae, CheolWon Na, Jaehoon Lee, Yumin Lee 等ACL 2026 · 被引用 1 次
