Does data sampling improve deep learning-based vulnerability detection? Yeas! and Nays!
Xu Yang, Shaowei Wang, Yi Li, Shaohua Wang
Abstract
Recent progress in Deep Learning (DL) has sparked interest in using DL to detect software vulnerabilities automatically and it has been demonstrated promising results at detecting vulnerabilities. However, one prominent and practical issue for vulnerability detection is data imbalance. Prior study observed that the performance of state-of-the-art (SOTA) DL-based vulnerability detection (DLVD) approaches drops precipitously in real world imbalanced data and a 73% drop of F1-score on average across studied approaches. Such a significant performance drop can disable the practical usage of any DLVD approaches. Data sampling is effective in alleviating data imbalance for machine learning models and has been demonstrated in various software engineering tasks. Therefore, in this study, we conducted a systematical and extensive study to assess the impact of data sampling for data imbalance problem in DLVD from two aspects: i) the effectiveness of DLVD, and ii) the ability of DLVD to reason correctly (making a decision based on real vulnerable statements). We found that in general, oversampling outperforms undersampling, and sampling on raw data outperforms sampling on latent space, typically random oversampling on raw data performs the best among all studied ones (including advanced one SMOTE and OSS). Surprisingly, OSS does not help alleviate the data imbalance issue in DLVD. If the recall is pursued, random undersampling is the best choice. Random oversampling on raw data also improves the ability of DLVD approaches for learning real vulnerable patterns. However, for a significant portion of cases (at least 33% in our datasets), DVLD approach cannot reason their prediction based on real vulnerable statements. We provide actionable suggestions and a roadmap to practitioners and researchers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5be39ed6-e46d-48cf-b94a-3d8d67115ebfCited by top-tier papers5
- SelfPiCo: Self-Guided Partial Code Execution with LLMsZhipeng Xue, Zhipeng Gao, Shaohua Wang, Xing Hu et al.ISSTA 2024 · 8 citations
- One-for-All Does Not Work! Enhancing Vulnerability Detection by Mixture-of-Experts (MoE)Xu Yang, Shaowei Wang, Jiayuan Zhou, Wenhan ZhuFSE 2025 · 7 citations
- Code Change Intention, Development Artifact, and History Vulnerability: Putting Them Together for Vulnerability Fix Detection by LLMXu Yang, Wenhan Zhu, Michael Pacheco, Jiayuan Zhou et al.FSE 2025 · 5 citations
- Three Heads Are Better Than One: A Multi-perspective Reasoning Framework for Enhanced Vulnerability DetectionXin Peng, Bo Lin, Jing Wang, Xiaoling Li et al.FSE 2026 · 1 citation
- Understanding Model Weaknesses: A Path to Strengthening DNN-Based Android Malware DetectionHaodong Li, Xiao Cheng, Yanjie Zhao, Guosheng Xu et al.ISSTA 2025
Builds on4
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 283 citations
- Finding A Needle in a Haystack: Automated Mining of Silent Vulnerability FixesJiayuan Zhou, Michael Pacheco, Zhiyuan Wan, Xin Xia et al.ASE 2021 · 84 citations
- PyExplainer: Explaining the Predictions of Just-In-Time Defect ModelsChanathip Pornprasit, Chakkrit Tantithamthavorn, Jirayus Jiarpakdee, Michael Fu et al.ASE 2021 · 52 citations
- VulDeePecker: A Deep Learning-Based System for Vulnerability DetectionZhen Li, Deqing Zou, Shouhuai Xu, Xinyu Ou et al.NDSS 2018
Related papers
- GVI: Guided Vulnerability Imagination for Boosting Deep Vulnerability DetectorsHeng Yong, Zhong Li, Minxue Pan, Tian Zhang et al.ICSE 2025 · 2 citations
- Distinguishing Look-Alike Innocent and Vulnerable Code by Subtle Semantic Representation Learning and ExplanationChao Ni, Xin Yin, Kaiwen Yang, Dehai Zhao et al.FSE 2023 · 42 citations
- VulSim: Leveraging Similarity of Multi-Dimensional Neighbor Embeddings for Vulnerability DetectionSamiha Shimmi, Ashiqur Rahman, Mohan Gadde, Hamed Okhravi et al.USENIX Security 2024 · 13 citations
- VULGEN: Realistic Vulnerability Generation Via Pattern Mining and Deep LearningYu Nong, Yuzhe Ou, Michael Pradel, Feng Chen et al.ICSE 2023 · 32 citations
- Generating realistic vulnerabilities via neural code editing: an empirical studyYu Nong, Yuzhe Ou, Michael Pradel, Feng Chen et al.FSE 2022 · 23 citations
