Silence of Commit Messages: An Empirical Study for Vulnerability Commit Message Generation using Large Language Models
Hao Shen, Ming Hu, Jiaye Li, Xiaofei Xie, Mingsong Chen
Abstract
The vulnerability commit message serves as crucial metadata for maintaining software within version control systems. Nonetheless, manually crafted vulnerability commit messages often lack detail or exhibit inconsistent formatting. Recently, the growing use of Large Language Models (LLMs) for code and natural language comprehension has opened avenues to automate the crafting of these messages. This paper systematically and thoroughly explores the generation of security patch commit messages in the context of LLMs, delving into topics such as dataset construction, evaluation method design, and the relationship between vulnerability types and submission structure. First, we explore the elements of commit messages using LLMs and integrate a questionnaire survey to pinpoint four essential types of information: summary, background, impact, and fix, which are essential for developers. This aims to establish a structured dataset of bug submissions and assess its quality. Next, we examine general automated evaluation techniques for assessing LLM-generated commit messages and find that GPT-3.5's evaluation methods align more closely with human judgment. Then, we conduct an organized investigation into how LLM generation effects vary across three principal vulnerability types, uncovering that LLMs' adaptability differs across vulnerabilities. Furthermore, we perform an exhaustive examination of generation quality across various components and find that LLMs excel at generating summaries but struggle to produce impact details. In particular, the smallest DeepSeek-Coder shows a semantic retention advantage in crafting backgrounds, whereas DeepSeek-V3 struggles with impact aspects. Lastly, we investigate the effects of different prompting strategies (e.g., zero-shot, few-shot prompts) and parameter settings (e.g., temperature and top_p) on the quality of commit message generation, finding that prompt and parameter configurations critically influence output quality, with model sensitivity varying.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Evaluating Generated Commit Messages with Large Language ModelsQunhong Zeng, Yuxia Zhang, Zexiong Ma, Bo Jiang et al.ICSE 2026
- Context Conquers Parameters: Outperforming Proprietary Llm in Commit Message GenerationAaron Imani, Iftekhar Ahmed, Mohammad MoshirpourICSE 2025 · 1 citation
- Only diff Is Not Enough: Generating Commit Messages Leveraging Reasoning and Action of Large Language ModelJiawei Li, David Faragó, Christian Petrov, Iftekhar AhmedFSE 2024 · 17 citations
- LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and BenchmarksSaad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce et al.S&P 2024 · 167 citations
- LLMBisect: Breaking Barriers in Bug Bisection with A Comparative Analysis PipelineZheng Zhang, Haonan Li, Xingyu Li, Hang Zhang et al.NDSS 2026 · 1 citation
