Lune

ISSTA2026Top-tier venue

Silence of Commit Messages: An Empirical Study for Vulnerability Commit Message Generation using Large Language Models

Hao Shen, Ming Hu, Jiaye Li, Xiaofei Xie, Mingsong Chen

2026Year

Abstract

The vulnerability commit message serves as crucial metadata for maintaining software within version control systems. Nonetheless, manually crafted vulnerability commit messages often lack detail or exhibit inconsistent formatting. Recently, the growing use of Large Language Models (LLMs) for code and natural language comprehension has opened avenues to automate the crafting of these messages. This paper systematically and thoroughly explores the generation of security patch commit messages in the context of LLMs, delving into topics such as dataset construction, evaluation method design, and the relationship between vulnerability types and submission structure. First, we explore the elements of commit messages using LLMs and integrate a questionnaire survey to pinpoint four essential types of information: summary, background, impact, and fix, which are essential for developers. This aims to establish a structured dataset of bug submissions and assess its quality. Next, we examine general automated evaluation techniques for assessing LLM-generated commit messages and find that GPT-3.5's evaluation methods align more closely with human judgment. Then, we conduct an organized investigation into how LLM generation effects vary across three principal vulnerability types, uncovering that LLMs' adaptability differs across vulnerabilities. Furthermore, we perform an exhaustive examination of generation quality across various components and find that LLMs excel at generating summaries but struggle to produce impact details. In particular, the smallest DeepSeek-Coder shows a semantic retention advantage in crafting backgrounds, whereas DeepSeek-V3 struggles with impact aspects. Lastly, we investigate the effects of different prompting strategies (e.g., zero-shot, few-shot prompts) and parameter settings (e.g., temperature and top_p) on the quality of commit message generation, finding that prompt and parameter configurations critically influence output quality, with model sensitivity varying.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines