Toward Improved Deep Learning-based Vulnerability Detection
Adriana Sejfia, Satyaki Das, Saad Shafiq, Nenad Medvidovic
摘要
Deep learning (DL) has been a common thread across several recent techniques for vulnerability detection. The rise of large, publicly available datasets of vulnerabilities has fueled the learning process underpinning these techniques. While these datasets help the DL-based vulnerability detectors, they also constrain these detectors' predictive abilities. Vulnerabilities in these datasets have to be represented in a certain way, e.g., code lines, functions, or program slices within which the vulnerabilities exist. We refer to this representation as a base unit. The detectors learn how base units can be vulnerable and then predict whether other base units are vulnerable. We have hypothesized that this focus on individual base units harms the ability of the detectors to properly detect those vulnerabilities that span multiple base units (or MBU vulnerabilities). For vulnerabilities such as these, a correct detection occurs when all comprising base units are detected as vulnerable. Verifying how existing techniques perform in detecting all parts of a vulnerability is important to establish their effectiveness for other downstream tasks. To evaluate our hypothesis, we conducted a study focusing on three prominent DL-based detectors: ReVeal, DeepWukong, and LineVul. Our study shows that all three detectors contain MBU vulnerabilities in their respective datasets. Further, we observed significant accuracy drops when detecting these types of vulnerabilities. We present our study and a framework that can be used to help DL-based detectors toward the proper inclusion of MBU vulnerabilities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller 等NDSS 2026 · 被引用 17 次
- Top Score on the Wrong Exam: On Benchmarking in Machine Learning for Vulnerability DetectionNiklas Risse, Jing Liu, Marcel BöhmeISSTA 2025 · 被引用 8 次
- Closing the Gap: A User Study on the Real-world Usefulness of AI-powered Vulnerability Detection & Repair in the IDEBenjamin Steenhoek, Kalpathy Sivaraman, Renata Saldivar Gonzalez, Yevhen Mohylevskyy 等ICSE 2025 · 被引用 3 次
- Are We Learning the Right Features? A Framework for Evaluating DL-Based Software Vulnerability Detection SolutionsSatyaki Das, Syeda Tasnim Fabiha, Saad Shafiq, Nenad MedvidovicICSE 2025 · 被引用 1 次
- Towards More Trustworthy Deep Code Models by Enabling Out-of-Distribution DetectionYanfu Yan, Viet Duong, Huajie Shao, Denys PoshyvanykICSE 2025 · 被引用 1 次
它引用的顶会 Paper3
- Data Quality for Software Vulnerability DatasetsRoland Croft, Muhammad Ali Babar, M. Mehdi KholoosiICSE 2023 · 被引用 138 次
- Identifying casualty changes in software patchesAdriana Sejfia, Yixue Zhao, Nenad MedvidovicFSE 2021 · 被引用 7 次
- VulDeePecker: A Deep Learning-Based System for Vulnerability DetectionZhen Li, Deqing Zou, Shouhuai Xu, Xinyu Ou 等NDSS 2018
相关 Paper
- An Empirical Study of Deep Learning Models for Vulnerability DetectionBenjamin Steenhoek, Md Mahbubur Rahman, Richard Jiles, Wei LeICSE 2023 · 被引用 107 次
- On the Effectiveness of Function-Level Vulnerability Detectors for Inter-Procedural VulnerabilitiesZhen Li, Ning Wang, Deqing Zou, Yating Li 等ICSE 2024 · 被引用 18 次
- Understanding and Tackling Label Errors in Deep Learning-Based Vulnerability Detection (Experience Paper)Xu Nie, Ningke Li, Kailong Wang, Shangguang Wang 等ISSTA 2023 · 被引用 23 次
- VulChecker: Graph-based Vulnerability Localization in Source CodeYisroel Mirsky, George Macon, Michael D. Brown, Carter Yagemann 等USENIX Security 2023
- GVI: Guided Vulnerability Imagination for Boosting Deep Vulnerability DetectorsHeng Yong, Zhong Li, Minxue Pan, Tian Zhang 等ICSE 2025 · 被引用 2 次
