Lune

ICSE2026Top-tier venue

Exploring and Improving Real-World Vulnerability Data Generation via Prompting Large Language Models

Guangbei Yi, Yu Nong, Minzhang Li, Haipeng Cai

2026Year

Abstract

Data-driven approaches have proven to be promising for vulnerability analysis, contingent on quality and sizable training data being available. Several dedicated vulnerability data generation techniques have demonstrated strong merits, yet they are limited to simple (single-line injection induced) vulnerabilities only and suffer from overfitting to seed samples. Large language models (LLMs) may overcome these fundamental limitations as they are known to be effective at generative tasks. However, it remains unclear how they would perform on the task of vulnerable sample generation. In this paper, we explore the potential and gaps of seven state-of-the-art (SOTA) LLMs for that task via prompting.

We reveal that the LLMs are capable of injecting vulnerabilities, with advanced prompting strategies such as few-shot in-context learning and our new vulnerability-introducing code-change semantics (VICS) guided prompting boosting the effectiveness, achieving up to 93% success rate on a synthetic dataset and 88% on realworld code. The LLMs can effectively perform both single-and multi-line injections, addressing a key limitation of prior work. Notably, they exhibit a strong preference for replacement edits, different from ground-truth patterns, and their effectiveness varies across CWE types. Furthermore, LLMs, particularly with VICS, outperform existing SOTA vulnerability generators with success rate improvements of up to 210%-343%. Crucially, the LLM-generated data substantially improves the performance of downstream DLbased vulnerability analysis models, especially with multi-line injections, boosting their accuracy by up to 70.1%. The generated samples also enhance the effectiveness of other LLMs for vulnerability analysis via RAG, with multi-line-injected samples yielding up to 16.5% gains. Most importantly, our findings reveal that, augmenting existing DL models with high-quality, LLM-generated data can lead to vulnerability analysis performance (up to 67.50% accuracy) superior to that of even the most advanced LLMs performing the same analysis (up to 27.40%), indicating the usefulness of the LLM-generated vulnerability data at present and in the longer term.

  • The author participated in this work primarily as an REU student.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ac485c0a-3e23-4362-b080-b25e4a6fb042

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines