Data Quality for Software Vulnerability Datasets
Roland Croft, Muhammad Ali Babar, M. Mehdi Kholoosi
Abstract
The use of learning-based techniques to achieve automated software vulnerability detection has been of longstanding interest within the software security domain. These data-driven solutions are enabled by large software vulnerability datasets used for training and benchmarking. However, we observe that the quality of the data powering these solutions is currently ill-considered, hindering the reliability and value of produced outcomes. Whilst awareness of software vulnerability data preparation challenges is growing, there has been little investigation into the potential negative impacts of software vulnerability data quality. For instance, we lack confirmation that vulnerability labels are correct or consistent. Our study seeks to address such shortcomings by inspecting five inherent data quality attributes for four state-of-the-art software vulnerability datasets and the subsequent impacts that issues can have on software vulnerability prediction models. Surprisingly, we found that all the analyzed datasets exhibit some data quality problems. In particular, we found 20–71% of vulnerability labels to be inaccurate in real-world datasets, and 17-99% of data points were duplicated. We observed that these issues could cause significant impacts on downstream models, either preventing effective model training or inflating benchmark performance. We advocate for the need to overcome such challenges. Our findings will enable better consideration and assessment of software vulnerability data quality in the future.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d01aa6e7-d09a-49db-9d07-52bb60aff6c1Cited by top-tier papers26
- Large Language Models for Code: Security Hardening and Adversarial TestingJingxuan He, Martin T. VechevCCS 2023 · 98 citations
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 86 citations
- Instruction Tuning for Secure Code GenerationJingxuan He, Mark Vero, Gabriela Krasnopolska, Martin T. VechevICML 2024 · 69 citations
- Vulnerability Detection with Code Language Models: How Far are We?Yangruibo Ding, Yanjun Fu, Omniyyah Ibrahim, Chawin Sitawarin et al.ICSE 2025 · 44 citations
- Coca: Improving and Explaining Graph Neural Network-Based Vulnerability Detection SystemsSicong Cao, Xiaobing Sun, Xiaoxue Wu, David Lo et al.ICSE 2024 · 27 citations
Builds on10
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 283 citations
- Finding A Needle in a Haystack: Automated Mining of Silent Vulnerability FixesJiayuan Zhou, Michael Pacheco, Zhiyuan Wan, Xin Xia et al.ASE 2021 · 84 citations
- On the Importance of Building High-quality Training Datasets for Neural Code SearchZhensu Sun, Li Li, Yan Liu, Xiaoning Du et al.ICSE 2022 · 67 citations
- Are we building on the rock? on the importance of data preprocessing for code summarizationLin Shi, Fangwen Mu, Xiao Chen, Song Wang et al.FSE 2022 · 65 citations
- Detecting False Alarms from Automatic Static Analysis Tools: How Far are We?Hong Jin Kang, Khai Loong Aw, David LoICSE 2022 · 42 citations
Related papers
- Understanding and Tackling Label Errors in Deep Learning-Based Vulnerability Detection (Experience Paper)Xu Nie, Ningke Li, Kailong Wang, Shangguang Wang et al.ISSTA 2023 · 23 citations
- Uncovering the Limits of Machine Learning for Automatic Vulnerability DetectionNiklas Risse, Marcel BöhmeUSENIX Security 2024 · 63 citations
- Out of Distribution, Out of Luck: How Well Can LLMs Trained on Vulnerability Datasets Detect Top 25 CWE Weaknesses?Yikun Li, Ngoc Tan Bui, Ting Zhang, Chengran Yang et al.ICSE 2026 · 2 citations
- An Empirical Study of Deep Learning Models for Vulnerability DetectionBenjamin Steenhoek, Md Mahbubur Rahman, Richard Jiles, Wei LeICSE 2023 · 107 citations
- Are We Learning the Right Features? A Framework for Evaluating DL-Based Software Vulnerability Detection SolutionsSatyaki Das, Syeda Tasnim Fabiha, Saad Shafiq, Nenad MedvidovicICSE 2025 · 1 citation
