Evaluating Code Summarization Techniques: A New Metric and an Empirical Characterization
Antonio Mastropaolo, Matteo Ciniselli, Massimiliano Di Penta, Gabriele Bavota
Abstract
Several code summarization techniques have been proposed in the literature to automatically document a code snippet or a function. Ideally, software developers should be involved in assessing the quality of the generated summaries. However, in most cases, researchers rely on automatic evaluation metrics such as BLEU, ROUGE, and METEOR. These metrics are all based on the same assumption: The higher the textual similarity between the generated summary and a reference summary written by developers, the higher its quality. However, there are two reasons for which this assumption falls short: (i) reference summaries, e.g., code comments collected by mining software repositories, may be of low quality or even outdated; (ii) generated summaries, while using a different wording than a reference one, could be semantically equivalent to it, thus still being suitable to document the code snippet. In this paper, we perform a thorough empirical investigation on the complementarity of different types of metrics in capturing the quality of a generated summary. Also, we propose to address the limitations of existing metrics by considering a new dimension, capturing the extent to which the generated summary aligns with the semantics of the documented code snippet, independently from the reference summary. To this end, we present a new metric based on contrastive learning to capture said aspect. We empirically show that the inclusion of this novel dimension enables a more effective representation of developers' evaluations regarding the quality of automatically generated summaries. CCS CONCEPTS • Software and its engineering → Documentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Source Code Summarization in the Era of Large Language ModelsWeisong Sun, Yun Miao, Yuekang Li, Hongyu Zhang et al.ICSE 2025 · 37 citations
- Calibration of Large Language Models on Code SummarizationYuvraj Virk, Premkumar T. Devanbu, Toufique AhmedFSE 2025 · 5 citations
- Calico: Automated Knowledge Calibration and Diagnosis for Elevating AI Mastery in Code TasksYuxin Qiu, Jie Hu, Qian Zhang, Heng YinISSTA 2024 · 3 citations
- ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code SummarizationSuyoung Bae, CheolWon Na, Jaehoon Lee, Yumin Lee et al.ACL 2026 · 1 citation
- SE-Jury: An LLM-as-Ensemble-Judge Metric for Narrowing the Gap with Human Evaluation in SEXin Zhou, Kisub Kim, Ting Zhang, Martin Weyssow et al.ASE 2025 · 1 citation
Builds on10
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
Related papers
- Reassessing automatic evaluation metrics for code summarization tasksDevjeet Roy, Sarah Fakhoury, Venera ArnaoudovaFSE 2021 · 103 citations
- Unsupervised Reference-Free Summary Quality Evaluation via Contrastive LearningHanlu Wu, Tengfei Ma, Lingfei Wu, Tariro Manyumwa et al.EMNLP 2020 · 47 citations
- SimLLM: Calculating Semantic Similarity in Code Summaries using a Large Language Model-Based ApproachXin Jin, Zhiqiang LinFSE 2024 · 8 citations
- Self-Supervised Contrastive Learning for Code Retrieval and Summarization via Semantic-Preserving TransformationsNghi D. Q. Bui, Yijun Yu, Lingxiao JiangSIGIR 2021 · 98 citations
- On the Evaluation of Neural Code SummarizationEnsheng Shi, Yanlin Wang, Lun Du, Junjie Chen et al.ICSE 2022 · 76 citations
