Exploring and Improving Real-World Vulnerability Data Generation via Prompting Large Language Models
Guangbei Yi, Yu Nong, Minzhang Li, Haipeng Cai
摘要
Data-driven approaches have proven to be promising for vulnerability analysis, contingent on quality and sizable training data being available. Several dedicated vulnerability data generation techniques have demonstrated strong merits, yet they are limited to simple (single-line injection induced) vulnerabilities only and suffer from overfitting to seed samples. Large language models (LLMs) may overcome these fundamental limitations as they are known to be effective at generative tasks. However, it remains unclear how they would perform on the task of vulnerable sample generation. In this paper, we explore the potential and gaps of seven state-of-the-art (SOTA) LLMs for that task via prompting.
We reveal that the LLMs are capable of injecting vulnerabilities, with advanced prompting strategies such as few-shot in-context learning and our new vulnerability-introducing code-change semantics (VICS) guided prompting boosting the effectiveness, achieving up to 93% success rate on a synthetic dataset and 88% on realworld code. The LLMs can effectively perform both single-and multi-line injections, addressing a key limitation of prior work. Notably, they exhibit a strong preference for replacement edits, different from ground-truth patterns, and their effectiveness varies across CWE types. Furthermore, LLMs, particularly with VICS, outperform existing SOTA vulnerability generators with success rate improvements of up to 210%-343%. Crucially, the LLM-generated data substantially improves the performance of downstream DLbased vulnerability analysis models, especially with multi-line injections, boosting their accuracy by up to 70.1%. The generated samples also enhance the effectiveness of other LLMs for vulnerability analysis via RAG, with multi-line-injected samples yielding up to 16.5% gains. Most importantly, our findings reveal that, augmenting existing DL models with high-quality, LLM-generated data can lead to vulnerability analysis performance (up to 67.50% accuracy) superior to that of even the most advanced LLMs performing the same analysis (up to 27.40%), indicating the usefulness of the LLM-generated vulnerability data at present and in the longer term.
- The author participated in this work primarily as an REU student.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 被引用 321 次
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 被引用 283 次
- Automated Repair of Programs from Large Language ModelsZhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury 等ICSE 2023 · 被引用 213 次
- VulRepair: a T5-based automated software vulnerability repairMichael Fu, Chakkrit Tantithamthavorn, Trung Le, Van Nguyen 等FSE 2022 · 被引用 206 次
- Repair Is Nearly Generation: Multilingual Program Repair with LLMsHarshit Joshi, José Pablo Cambronero Sánchez, Sumit Gulwani, Vu Le 等AAAI 2023 · 被引用 182 次
相关 Paper
- Generating realistic vulnerabilities via neural code editing: an empirical studyYu Nong, Yuzhe Ou, Michael Pradel, Feng Chen 等FSE 2022 · 被引用 23 次
- Hit The Bullseye On The First Shot: Improving LLMs Using Multi-Sample Self-Reward Feedback for Vulnerability RepairRui Jiao, Yue Zhang, Jinku Li, Jianfeng MaASE 2025
- GVI: Guided Vulnerability Imagination for Boosting Deep Vulnerability DetectorsHeng Yong, Zhong Li, Minxue Pan, Tian Zhang 等ICSE 2025 · 被引用 2 次
- Enhancing Vulnerability Detection via Inter-procedural Semantic CompletionBozhi Wu, Chengjie Liu, Zhiming Li, Yushi Cao 等ISSTA 2025 · 被引用 2 次
- Repairing LLM Executions for Secure Automatic ProgrammingAli El Husseini, Yacine Izza, Blaise Genest, Abhik RoychoudhuryICSE 2026
