Data Quality for Software Vulnerability Datasets
Roland Croft, Muhammad Ali Babar, M. Mehdi Kholoosi
摘要
The use of learning-based techniques to achieve automated software vulnerability detection has been of longstanding interest within the software security domain. These data-driven solutions are enabled by large software vulnerability datasets used for training and benchmarking. However, we observe that the quality of the data powering these solutions is currently ill-considered, hindering the reliability and value of produced outcomes. Whilst awareness of software vulnerability data preparation challenges is growing, there has been little investigation into the potential negative impacts of software vulnerability data quality. For instance, we lack confirmation that vulnerability labels are correct or consistent. Our study seeks to address such shortcomings by inspecting five inherent data quality attributes for four state-of-the-art software vulnerability datasets and the subsequent impacts that issues can have on software vulnerability prediction models. Surprisingly, we found that all the analyzed datasets exhibit some data quality problems. In particular, we found 20–71% of vulnerability labels to be inaccurate in real-world datasets, and 17-99% of data points were duplicated. We observed that these issues could cause significant impacts on downstream models, either preventing effective model training or inflating benchmark performance. We advocate for the need to overcome such challenges. Our findings will enable better consideration and assessment of software vulnerability data quality in the future.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Large Language Models for Code: Security Hardening and Adversarial TestingJingxuan He, Martin T. VechevCCS 2023 · 被引用 98 次
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 被引用 86 次
- Instruction Tuning for Secure Code GenerationJingxuan He, Mark Vero, Gabriela Krasnopolska, Martin T. VechevICML 2024 · 被引用 69 次
- Vulnerability Detection with Code Language Models: How Far are We?Yangruibo Ding, Yanjun Fu, Omniyyah Ibrahim, Chawin Sitawarin 等ICSE 2025 · 被引用 44 次
- Coca: Improving and Explaining Graph Neural Network-Based Vulnerability Detection SystemsSicong Cao, Xiaobing Sun, Xiaoxue Wu, David Lo 等ICSE 2024 · 被引用 27 次
它引用的顶会 Paper10
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 被引用 283 次
- Finding A Needle in a Haystack: Automated Mining of Silent Vulnerability FixesJiayuan Zhou, Michael Pacheco, Zhiyuan Wan, Xin Xia 等ASE 2021 · 被引用 84 次
- On the Importance of Building High-quality Training Datasets for Neural Code SearchZhensu Sun, Li Li, Yan Liu, Xiaoning Du 等ICSE 2022 · 被引用 67 次
- Are we building on the rock? on the importance of data preprocessing for code summarizationLin Shi, Fangwen Mu, Xiao Chen, Song Wang 等FSE 2022 · 被引用 65 次
- Detecting False Alarms from Automatic Static Analysis Tools: How Far are We?Hong Jin Kang, Khai Loong Aw, David LoICSE 2022 · 被引用 42 次
相关 Paper
- Understanding and Tackling Label Errors in Deep Learning-Based Vulnerability Detection (Experience Paper)Xu Nie, Ningke Li, Kailong Wang, Shangguang Wang 等ISSTA 2023 · 被引用 23 次
- Uncovering the Limits of Machine Learning for Automatic Vulnerability DetectionNiklas Risse, Marcel BöhmeUSENIX Security 2024 · 被引用 63 次
- Out of Distribution, Out of Luck: How Well Can LLMs Trained on Vulnerability Datasets Detect Top 25 CWE Weaknesses?Yikun Li, Ngoc Tan Bui, Ting Zhang, Chengran Yang 等ICSE 2026 · 被引用 2 次
- An Empirical Study of Deep Learning Models for Vulnerability DetectionBenjamin Steenhoek, Md Mahbubur Rahman, Richard Jiles, Wei LeICSE 2023 · 被引用 107 次
- Are We Learning the Right Features? A Framework for Evaluating DL-Based Software Vulnerability Detection SolutionsSatyaki Das, Syeda Tasnim Fabiha, Saad Shafiq, Nenad MedvidovicICSE 2025 · 被引用 1 次
