Are We Learning the Right Features? A Framework for Evaluating DL-Based Software Vulnerability Detection Solutions
Satyaki Das, Syeda Tasnim Fabiha, Saad Shafiq, Nenad Medvidovic
摘要
Recent research has revealed that the reported results of an emerging body of deep learning-based techniques for detecting software vulnerabilities are not reproducible, either across different datasets or on unseen samples. This paper aims to provide the foundation for properly evaluating the research in this domain. We do so by analyzing prior work and existing vulnerability datasets for the syntactic and semantic features of code that contribute to vulnerability, as well as features that falsely correlate with vulnerability. We provide a novel, uniform representation to capture both sets of features, and use this representation to detect the presence of both vulnerability and spurious features in code. To this end, we design two types of code perturbations: feature preserving perturbations (FPP) ensure that the vulnerability feature remains in a given code sample, while feature eliminating perturbations (FEP) eliminate the feature from the code sample. These perturbations aim to measure the influence of spurious and vulnerability features on the predictions of a given vulnerability detection solution. To evaluate how the two classes of perturbations influence predictions, we conducted a large-scale empirical study on five state-of-the-art DL-based vulnerability detectors. Our study shows that, for vulnerability features, only <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> of FPPs yield the undesirable effect of a prediction changing among the five detectors on average. However, on average, <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> of FEPs yield the undesirable effect of retaining the vulnerability predictions. For spurious features, we observed that FPPs yielded a drop in recall up to 29 % for graph-based detectors. We present the reasons underlying these results and suggest strategies for improving DNN-based vulnerability detectors. We provide our perturbation-based evaluation framework as a public resource to enable independent future evaluation of vulnerability detectors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- VUDDY: A Scalable Approach for Vulnerable Code Clone DiscoverySeulbae Kim, Seunghoon Woo, Heejo Lee, Hakjoo OhS&P 2017 · 被引用 388 次
- Natural Attack for Pre-trained Models of CodeZhou Yang, Jieke Shi, Junda He, David LoICSE 2022 · 被引用 150 次
- Dataflow Analysis-Inspired Deep Learning for Efficient Vulnerability DetectionBenjamin Steenhoek, Hongyang Gao, Wei LeICSE 2024 · 被引用 54 次
- DeepVD: Toward Class-Separation Features for Neural Network Vulnerability DetectionWenbo Wang, Tien N. Nguyen, Shaohua Wang, Yi Li 等ICSE 2023 · 被引用 32 次
- Probing model signal-awareness via prediction-preserving input minimizationSahil Suneja, Yunhui Zheng, Yufan Zhuang, Jim Alain Laredo 等FSE 2021 · 被引用 29 次
相关 Paper
- VULGEN: Realistic Vulnerability Generation Via Pattern Mining and Deep LearningYu Nong, Yuzhe Ou, Michael Pradel, Feng Chen 等ICSE 2023 · 被引用 32 次
- Toward Improved Deep Learning-based Vulnerability DetectionAdriana Sejfia, Satyaki Das, Saad Shafiq, Nenad MedvidovicICSE 2024 · 被引用 14 次
- Uncovering the Limits of Machine Learning for Automatic Vulnerability DetectionNiklas Risse, Marcel BöhmeUSENIX Security 2024 · 被引用 63 次
- VulSim: Leveraging Similarity of Multi-Dimensional Neighbor Embeddings for Vulnerability DetectionSamiha Shimmi, Ashiqur Rahman, Mohan Gadde, Hamed Okhravi 等USENIX Security 2024 · 被引用 13 次
- An Empirical Study of Deep Learning Models for Vulnerability DetectionBenjamin Steenhoek, Md Mahbubur Rahman, Richard Jiles, Wei LeICSE 2023 · 被引用 107 次
