Evaluating Generated Commit Messages with Large Language Models
Qunhong Zeng, Yuxia Zhang, Zexiong Ma, Bo Jiang, Ningyuan Sun, Klaas-Jan Stol, Xingyu Mou, Hui Liu
Abstract
Commit messages are essential in software development as they serve to document and explain code changes. Yet, their quality often falls short in practice, with studies showing significant proportions of empty or inadequate messages. While automated commit message generation has advanced significantly, particularly with Large Language Models (LLMs), the evaluation of generated messages remains challenging. Traditional reference-based automatic metrics like BLEU, ROUGE-L, and METEOR have notable limitations in assessing commit message quality, as they assume a one-to-one mapping between code changes and commit messages, leading researchers to rely on resource-intensive human evaluation. This study investigates the potential of LLMs as automated evaluators for commit message quality. Through systematic experimentation with various prompt strategies and state-of-the-art LLMs, we demonstrate that LLMs combining Chain-of-Thought reasoning with few-shot demonstrations achieve near human-level evaluation proficiency. Our LLM-based evaluator significantly outperforms traditional metrics while maintaining acceptable reproducibility, robustness, and fairness levels despite some inherent variability. This work conducts a comprehensive preliminary study on using LLMs for commit message evaluation, offering a scalable alternative to human assessment while maintaining high-quality evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 465aa7bf-deaf-4de7-9cc0-7e6851d35d7aBuilds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- Only diff Is Not Enough: Generating Commit Messages Leveraging Reasoning and Action of Large Language ModelJiawei Li, David Faragó, Christian Petrov, Iftekhar AhmedFSE 2024 · 17 citations
- Context Conquers Parameters: Outperforming Proprietary Llm in Commit Message GenerationAaron Imani, Iftekhar Ahmed, Mohammad MoshirpourICSE 2025 · 1 citation
- Silence of Commit Messages: An Empirical Study for Vulnerability Commit Message Generation using Large Language ModelsHao Shen, Ming Hu, Jiaye Li, Xiaofei Xie et al.ISSTA 2026
- An Empirical Study on Commit Message Generation Using LLMs via In-Context LearningYifan Wu, Yunpeng Wang, Ying Li, Wei Tao et al.ICSE 2025 · 1 citation
- Revisiting Learning-based Commit Message GenerationJinhao Dong, Yiling Lou, Dan Hao, Lin TanICSE 2023 · 8 citations
