Path-sensitive code embedding via contrastive learning for software vulnerability detection
Xiao Cheng, Guanqin Zhang, Haoyu Wang, Yulei Sui
摘要
Machine learning and its promising branch deep learning have shown success in a wide range of application domains. Recently, much effort has been expended on applying deep learning techniques (e.g., graph neural networks) to static vulnerability detection as an alternative to conventional bug detection methods. To obtain the structural information of code, current learning approaches typically abstract a program in the form of graphs (e.g., data-flow graphs, abstract syntax trees), and then train an underlying classification model based on the (sub)graphs of safe and vulnerable code fragments for vulnerability prediction. However, these models are still insufficient for precise bug detection, because the objective of these models is to produce classification results rather than comprehending the semantics of vulnerabilities, e.g., pinpoint bug triggering paths, which are essential for static bug detection. This paper presents ContraFlow, a selective yet precise contrastive value-flow embedding approach to statically detect software vulnerabilities. The novelty of ContraFlow lies in selecting and preserving feasible value-flow (aka program dependence) paths through a pretrained path embedding model using self-supervised contrastive learning, thus significantly reducing the amount of labeled data required for training expensive downstream models for path-based vulnerability detection. We evaluated ContraFlow using 288 real-world projects by comparing eight recent learningbased approaches. ContraFlow outperforms these eight baselines by up to 334.1%, 317.9%, 58.3% for informedness, markedness and F1 Score, and ContraFlow achieves up to 450.0%, 192.3%, 450.0% improvement for mean statement recall, mean statement precision and mean IoU respectively in terms of locating buggy statements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Dataflow Analysis-Inspired Deep Learning for Efficient Vulnerability DetectionBenjamin Steenhoek, Hongyang Gao, Wei LeICSE 2024 · 被引用 54 次
- Distinguishing Look-Alike Innocent and Vulnerable Code by Subtle Semantic Representation Learning and ExplanationChao Ni, Xin Yin, Kaiwen Yang, Dehai Zhao 等FSE 2023 · 被引用 42 次
- Coca: Improving and Explaining Graph Neural Network-Based Vulnerability Detection SystemsSicong Cao, Xiaobing Sun, Xiaoxue Wu, David Lo 等ICSE 2024 · 被引用 27 次
- When Less is Enough: Positive and Unlabeled Learning Model for Vulnerability DetectionXin-Cheng Wen, Xinchen Wang, Cuiyun Gao, Shaohua Wang 等ASE 2023 · 被引用 20 次
- Combining Structured Static Code Information and Dynamic Symbolic Traces for Software Vulnerability PredictionHuanting Wang, Zhanyong Tang, Shin Hwei Tan, Jie Wang 等ICSE 2024 · 被引用 15 次
它引用的顶会 Paper12
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- A Security Analysis of HoneywordsDing Wang, Haibo Cheng, Ping Wang, Jeff Yan 等NDSS 2018 · 被引用 1,102 次
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 被引用 283 次
- Self-supervised Video Representation Learning Using Inter-intra Contrastive FrameworkLi Tao, Xueting Wang, Toshihiko YamasakiACM MM 2020 · 被引用 110 次
相关 Paper
- DeepVD: Toward Class-Separation Features for Neural Network Vulnerability DetectionWenbo Wang, Tien N. Nguyen, Shaohua Wang, Yi Li 等ICSE 2023 · 被引用 32 次
- Improving Smart Contract Security with Contrastive Learning-based Vulnerability DetectionYizhou Chen, Zeyu Sun, Zhihao Gong, Dan HaoICSE 2024 · 被引用 44 次
- Vulnerability Detection with Graph Simplification and Enhanced Graph Representation LearningXin-Cheng Wen, Yupan Chen, Cuiyun Gao, Hongyu Zhang 等ICSE 2023 · 被引用 77 次
- Learning to Locate and Describe VulnerabilitiesJian Zhang, Shangqing Liu, Xu Wang, Tianlin Li 等ASE 2023 · 被引用 8 次
- MVD: Memory-Related Vulnerability Detection Based on Flow-Sensitive Graph Neural NetworksSicong Cao, Xiaobing Sun, Lili Bo, Rongxin Wu 等ICSE 2022 · 被引用 100 次
