Revisiting Learning-based Commit Message Generation
Jinhao Dong, Yiling Lou, Dan Hao, Lin Tan
摘要
Commit messages summarize code changes and help developers understand the intention. To alleviate human efforts in writing commit messages, researchers have proposed various automated commit message generation techniques, among which learning-based techniques have achieved great success in recent years. However, existing evaluation on learning-based commit message generation relies on the automatic metrics (e.g., BLEU) widely used in natural language processing (NLP) tasks, which are aggregated scores calculated based on the similarity between generated commit messages and the ground truth. Therefore, it remains unclear what generated commit messages look like and what kind of commit messages could be precisely generated by existing learning-based techniques. To fill this knowledge gap, this work performs the first study to systematically investigate the detailed commit messages generated by learning-based techniques. In particular, we first investigate the frequent patterns of the commit messages generated by state-of-the-art learning-based techniques. Surprisingly, we find the majority ( 90%) of their generated commit messages belong to simple patterns (i.e., addition/removal/fix/avoidance patterns). To further explore the reasons, we then study the impact of datasets, input representations, and model components. We surprisingly find that existing learning-based techniques have competitive performance even when the inputs are only represented by change marks (i.e., “+”/“-”/“ ”), It indicates that existing learning-based techniques poorly utilize syntax and semantics in the code while mostly focusing on change marks, which could be the major reason for generating so many pattern-matching commit messages. We also find that the pattern ratio in the training set might also positively affect the pattern ratio of generated commit messages; and model components might have different impact on the pattern ratio.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Understanding Code Changes Practically with Small-Scale Language ModelsCong Li, Zhaogui Xu, Peng Di, Dongxia Wang 等ASE 2024 · 被引用 4 次
- Mutual Learning-Based Framework for Enhancing Robustness of Code Models via Adversarial TrainingYangsen Wang, Yizhou Chen, Yifan Zhao, Zhihao Gong 等ASE 2024 · 被引用 3 次
- Context Conquers Parameters: Outperforming Proprietary Llm in Commit Message GenerationAaron Imani, Iftekhar Ahmed, Mohammad MoshirpourICSE 2025 · 被引用 1 次
它引用的顶会 Paper3
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 被引用 169 次
- What Makes a Good Commit Message?Yingchen Tian, Yuxia Zhang, Klaas-Jan Stol, Lin Jiang 等ICSE 2022 · 被引用 90 次
- FIRA: Fine-Grained Graph-Based Code Change Representation for Automated Commit Message GenerationJinhao Dong, Yiling Lou, Qihao Zhu, Zeyu Sun 等ICSE 2022 · 被引用 50 次
相关 Paper
- An Empirical Study on Learning-based Techniques for Explicit and Implicit Commit Messages GenerationZhiquan Huang, Yuan Huang, Xiangping Chen, Xiaocong Zhou 等ASE 2024 · 被引用 2 次
- Evaluating Generated Commit Messages with Large Language ModelsQunhong Zeng, Yuxia Zhang, Zexiong Ma, Bo Jiang 等ICSE 2026
- COME: Commit Message Generation with Modification EmbeddingYichen He, Liran Wang, Kaiyi Wang, Yupeng Zhang 等ISSTA 2023 · 被引用 14 次
- RACE: Retrieval-augmented Commit Message GenerationEnsheng Shi, Yanlin Wang, Wei Tao, Lun Du 等EMNLP 2022 · 被引用 36 次
- An Empirical Study on Commit Message Generation Using LLMs via In-Context LearningYifan Wu, Yunpeng Wang, Ying Li, Wei Tao 等ICSE 2025 · 被引用 1 次
