Does data sampling improve deep learning-based vulnerability detection? Yeas! and Nays!
Xu Yang, Shaowei Wang, Yi Li, Shaohua Wang
摘要
Recent progress in Deep Learning (DL) has sparked interest in using DL to detect software vulnerabilities automatically and it has been demonstrated promising results at detecting vulnerabilities. However, one prominent and practical issue for vulnerability detection is data imbalance. Prior study observed that the performance of state-of-the-art (SOTA) DL-based vulnerability detection (DLVD) approaches drops precipitously in real world imbalanced data and a 73% drop of F1-score on average across studied approaches. Such a significant performance drop can disable the practical usage of any DLVD approaches. Data sampling is effective in alleviating data imbalance for machine learning models and has been demonstrated in various software engineering tasks. Therefore, in this study, we conducted a systematical and extensive study to assess the impact of data sampling for data imbalance problem in DLVD from two aspects: i) the effectiveness of DLVD, and ii) the ability of DLVD to reason correctly (making a decision based on real vulnerable statements). We found that in general, oversampling outperforms undersampling, and sampling on raw data outperforms sampling on latent space, typically random oversampling on raw data performs the best among all studied ones (including advanced one SMOTE and OSS). Surprisingly, OSS does not help alleviate the data imbalance issue in DLVD. If the recall is pursued, random undersampling is the best choice. Random oversampling on raw data also improves the ability of DLVD approaches for learning real vulnerable patterns. However, for a significant portion of cases (at least 33% in our datasets), DVLD approach cannot reason their prediction based on real vulnerable statements. We provide actionable suggestions and a roadmap to practitioners and researchers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SelfPiCo: Self-Guided Partial Code Execution with LLMsZhipeng Xue, Zhipeng Gao, Shaohua Wang, Xing Hu 等ISSTA 2024 · 被引用 8 次
- One-for-All Does Not Work! Enhancing Vulnerability Detection by Mixture-of-Experts (MoE)Xu Yang, Shaowei Wang, Jiayuan Zhou, Wenhan ZhuFSE 2025 · 被引用 7 次
- Code Change Intention, Development Artifact, and History Vulnerability: Putting Them Together for Vulnerability Fix Detection by LLMXu Yang, Wenhan Zhu, Michael Pacheco, Jiayuan Zhou 等FSE 2025 · 被引用 5 次
- Three Heads Are Better Than One: A Multi-perspective Reasoning Framework for Enhanced Vulnerability DetectionXin Peng, Bo Lin, Jing Wang, Xiaoling Li 等FSE 2026 · 被引用 1 次
- Understanding Model Weaknesses: A Path to Strengthening DNN-Based Android Malware DetectionHaodong Li, Xiao Cheng, Yanjie Zhao, Guosheng Xu 等ISSTA 2025
它引用的顶会 Paper4
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 被引用 283 次
- Finding A Needle in a Haystack: Automated Mining of Silent Vulnerability FixesJiayuan Zhou, Michael Pacheco, Zhiyuan Wan, Xin Xia 等ASE 2021 · 被引用 84 次
- PyExplainer: Explaining the Predictions of Just-In-Time Defect ModelsChanathip Pornprasit, Chakkrit Tantithamthavorn, Jirayus Jiarpakdee, Michael Fu 等ASE 2021 · 被引用 52 次
- VulDeePecker: A Deep Learning-Based System for Vulnerability DetectionZhen Li, Deqing Zou, Shouhuai Xu, Xinyu Ou 等NDSS 2018
相关 Paper
- GVI: Guided Vulnerability Imagination for Boosting Deep Vulnerability DetectorsHeng Yong, Zhong Li, Minxue Pan, Tian Zhang 等ICSE 2025 · 被引用 2 次
- Distinguishing Look-Alike Innocent and Vulnerable Code by Subtle Semantic Representation Learning and ExplanationChao Ni, Xin Yin, Kaiwen Yang, Dehai Zhao 等FSE 2023 · 被引用 42 次
- VulSim: Leveraging Similarity of Multi-Dimensional Neighbor Embeddings for Vulnerability DetectionSamiha Shimmi, Ashiqur Rahman, Mohan Gadde, Hamed Okhravi 等USENIX Security 2024 · 被引用 13 次
- VULGEN: Realistic Vulnerability Generation Via Pattern Mining and Deep LearningYu Nong, Yuzhe Ou, Michael Pradel, Feng Chen 等ICSE 2023 · 被引用 32 次
- Generating realistic vulnerabilities via neural code editing: an empirical studyYu Nong, Yuzhe Ou, Michael Pradel, Feng Chen 等FSE 2022 · 被引用 23 次
