VGX: Large-Scale Sample Generation for Boosting Learning-Based Software Vulnerability Analyses
Yu Nong, Richard Fang, Guangbei Yi, Kunsong Zhao, Xiapu Luo, Feng Chen, Haipeng Cai
摘要
Accompanying the successes of learning-based defensive software vulnerability analyses is the lack of large and quality sets of labeled vulnerable program samples, which impedes further advancement of those defenses. Existing automated sample generation approaches have shown potentials yet still fall short of practical expectations due to the high noise in the generated samples. This paper proposes VGX, a new technique aimed for large-scale generation of high-quality vulnerability datasets. Given a normal program, VGX identifies the code contexts in which vulnerabilities can be injected, using a customized Transformer featured with a new value-flowbased position encoding and pre-trained against new objectives particularly for learning code structure and context. Then, VGX materializes vulnerability-injection code editing in the identified contexts using patterns of such edits obtained from both historical fixes and human knowledge about real-world vulnerabilities. Compared to four state-of-the-art (SOTA) (i.e., pattern-, Transformer-, GNN-, and pattern+Transformer-based) baselines, VGX achieved 99.09-890.06% higher F1 and 22.45%-328.47% higher label accuracy. For in-the-wild sample production, VGX generated 150,392 vulnerable samples, from which we randomly chose 10% to assess how much these samples help vulnerability detection, localization, and repair. Our results show SOTA techniques for these three application tasks achieved 19.15-330.80% higher F1, 12.86-19.31% higher top-10 accuracy, and 85.02-99.30% higher top-50 accuracy, respectively, by adding those samples to their original training data. These samples also helped a SOTA vulnerability detector discover 13 more real-world vulnerabilities (CVEs) in critical systems (e.g., Linux kernel) that would be missed by the original model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Learning to Detect and Localize Multilingual BugsHaoran Yang, Yu Nong, Tao Zhang, Xiapu Luo 等FSE 2024 · 被引用 8 次
- SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability AnalysisYansong Li, Paula Branco, Alexander M. Hoole, Manish Marwah 等S&P 2025
- APPATCH: Automated Adaptive Prompting Large Language Models for Real-World Software Vulnerability PatchingYu Nong, Haoran Yang, Long Cheng, Hongxin Hu 等USENIX Security 2025
- NEXUS: Towards Accurate and Scalable Mapping between Vulnerabilities and Attack TechniquesEhsan Khodayarseresht, Suryadipta Majumdar, Serguei A. Mokhov, Mourad DebbabiNDSS 2026
- Exploring and Improving Real-World Vulnerability Data Generation via Prompting Large Language ModelsGuangbei Yi, Yu Nong, Minzhang Li, Haipeng CaiICSE 2026
它引用的顶会 Paper21
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 被引用 283 次
- Hoppity: Learning Graph Transformations to Detect and Fix Bugs in ProgramsElizabeth Dinella, Hanjun Dai, Ziyang Li, Mayur Naik 等ICLR 2020 · 被引用 212 次
- VulRepair: a T5-based automated software vulnerability repairMichael Fu, Chakkrit Tantithamthavorn, Trung Le, Van Nguyen 等FSE 2022 · 被引用 206 次
相关 Paper
- VULGEN: Realistic Vulnerability Generation Via Pattern Mining and Deep LearningYu Nong, Yuzhe Ou, Michael Pradel, Feng Chen 等ICSE 2023 · 被引用 32 次
- Generating realistic vulnerabilities via neural code editing: an empirical studyYu Nong, Yuzhe Ou, Michael Pradel, Feng Chen 等FSE 2022 · 被引用 23 次
- Dataflow Analysis-Inspired Deep Learning for Efficient Vulnerability DetectionBenjamin Steenhoek, Hongyang Gao, Wei LeICSE 2024 · 被引用 54 次
- Using Safety Properties to Generate Vulnerability PatchesZhen Huang, David Lie, Gang Tan, Trent JaegerS&P 2019 · 被引用 91 次
- CTX-Coder: Cross-Attention Architectures Empower LLMs for Long-Context Vulnerability DetectionJujie Wang, Kangfeng Zheng, Bin Wu, Chunhua Wu 等AAAI 2026
