Lune

ISSTA2026顶会

Silence of Commit Messages: An Empirical Study for Vulnerability Commit Message Generation using Large Language Models

Hao Shen, Ming Hu, Jiaye Li, Xiaofei Xie, Mingsong Chen

2026年份

摘要

The vulnerability commit message serves as crucial metadata for maintaining software within version control systems. Nonetheless, manually crafted vulnerability commit messages often lack detail or exhibit inconsistent formatting. Recently, the growing use of Large Language Models (LLMs) for code and natural language comprehension has opened avenues to automate the crafting of these messages. This paper systematically and thoroughly explores the generation of security patch commit messages in the context of LLMs, delving into topics such as dataset construction, evaluation method design, and the relationship between vulnerability types and submission structure. First, we explore the elements of commit messages using LLMs and integrate a questionnaire survey to pinpoint four essential types of information: summary, background, impact, and fix, which are essential for developers. This aims to establish a structured dataset of bug submissions and assess its quality. Next, we examine general automated evaluation techniques for assessing LLM-generated commit messages and find that GPT-3.5's evaluation methods align more closely with human judgment. Then, we conduct an organized investigation into how LLM generation effects vary across three principal vulnerability types, uncovering that LLMs' adaptability differs across vulnerabilities. Furthermore, we perform an exhaustive examination of generation quality across various components and find that LLMs excel at generating summaries but struggle to produce impact details. In particular, the smallest DeepSeek-Coder shows a semantic retention advantage in crafting backgrounds, whereas DeepSeek-V3 struggles with impact aspects. Lastly, we investigate the effects of different prompting strategies (e.g., zero-shot, few-shot prompts) and parameter settings (e.g., temperature and top_p) on the quality of commit message generation, finding that prompt and parameter configurations critically influence output quality, with model sensitivity varying.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖