VULGEN: Realistic Vulnerability Generation Via Pattern Mining and Deep Learning
Yu Nong, Yuzhe Ou, Michael Pradel, Feng Chen, Haipeng Cai
摘要
Building new, powerful data-driven defenses against prevalent software vulnerabilities needs sizable, quality vulnerability datasets, so does large-scale benchmarking of existing defense solutions. Automatic data generation would promisingly meet the need, yet there is little work aimed to generate much-needed quality vulnerable samples. Meanwhile, existing similar and adaptable techniques suffer critical limitations for that purpose. In this paper, we present VULGEN, the first injection-based vulnerability-generation technique that is not limited to a particular class of vulnerabilities. VULGEN combines the strengths of deterministic (pattern-based) and probabilistic (deep-learning/DL-based) program transformation approaches while mutually overcoming respective weaknesses. This is achieved through close collaborations between pattern mining/application and DL-based injection localization, which separates the concerns with how and where to inject. By leveraging large, pretrained programming language modeling and only learning locations, VULGEN mitigates its own needs for quality vulnerability data (for training the localization model). Extensive evaluations show that VULGEN significantly outperforms a state-of-the-art (SOTA) pattern-based peer technique as well as both Transformer- and GNN-based approaches in terms of the percentages of generated samples that are vulnerable and those also exactly matching the ground truth (by 38.0-430.1% and 16.3-158.2%, respectively). The VULGEN-generated samples led to substantial performance improvements for two SOTA DL-based vulnerability detectors (by up to 31.8% higher in F1), close to those brought by the ground-truth real-world samples and much higher than those by the same numbers of existing synthetic samples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- An Empirical Study on Fine-Tuning Large Language Models of Code for Automated Program RepairKai Huang, Xiangxin Meng, Jian Zhang, Yang Liu 等ASE 2023 · 被引用 91 次
- Coca: Improving and Explaining Graph Neural Network-Based Vulnerability Detection SystemsSicong Cao, Xiaobing Sun, Xiaoxue Wu, David Lo 等ICSE 2024 · 被引用 27 次
- VGX: Large-Scale Sample Generation for Boosting Learning-Based Software Vulnerability AnalysesYu Nong, Richard Fang, Guangbei Yi, Kunsong Zhao 等ICSE 2024 · 被引用 23 次
- Learning to Detect and Localize Multilingual BugsHaoran Yang, Yu Nong, Tao Zhang, Xiapu Luo 等FSE 2024 · 被引用 8 次
- Template-Guided Program Repair in the Era of Large Language ModelsKai Huang, Jian Zhang, Xiangxin Meng, Yang LiuICSE 2025 · 被引用 7 次
它引用的顶会 Paper13
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- CURE: Code-Aware Neural Machine Translation for Automatic Program RepairNan Jiang, Thibaud Lutellier, Lin TanICSE 2021 · 被引用 267 次
- Less training, more repairing please: revisiting automated program repair via zero-shot learningChunqiu Steven Xia, Lingming ZhangFSE 2022 · 被引用 223 次
- VulRepair: a T5-based automated software vulnerability repairMichael Fu, Chakkrit Tantithamthavorn, Trung Le, Van Nguyen 等FSE 2022 · 被引用 206 次
- TFix: Learning to Fix Coding Errors with a Text-to-Text TransformerBerkay Berabi, Jingxuan He, Veselin Raychev, Martin T. VechevICML 2021 · 被引用 143 次
相关 Paper
- Generating realistic vulnerabilities via neural code editing: an empirical studyYu Nong, Yuzhe Ou, Michael Pradel, Feng Chen 等FSE 2022 · 被引用 23 次
- Dataflow Analysis-Inspired Deep Learning for Efficient Vulnerability DetectionBenjamin Steenhoek, Hongyang Gao, Wei LeICSE 2024 · 被引用 54 次
- Are We Learning the Right Features? A Framework for Evaluating DL-Based Software Vulnerability Detection SolutionsSatyaki Das, Syeda Tasnim Fabiha, Saad Shafiq, Nenad MedvidovicICSE 2025 · 被引用 1 次
- GVI: Guided Vulnerability Imagination for Boosting Deep Vulnerability DetectorsHeng Yong, Zhong Li, Minxue Pan, Tian Zhang 等ICSE 2025 · 被引用 2 次
- Exploring and Improving Real-World Vulnerability Data Generation via Prompting Large Language ModelsGuangbei Yi, Yu Nong, Minzhang Li, Haipeng CaiICSE 2026
