Pre-training by Predicting Program Dependencies for Vulnerability Analysis Tasks
Zhongxin Liu, Zhijie Tang, Junwei Zhang, Xin Xia, Xiaohu Yang
摘要
Vulnerability analysis is crucial for software security. Inspired by the success of pre-trained models on software engineering tasks, this work focuses on using pre-training techniques to enhance the understanding of vulnerable code and boost vulnerability analysis. The code understanding ability of a pre-trained model is highly related to its pre-training objectives. The semantic structure, e.g., control and data dependencies, of code is important for vulnerability analysis. However, existing pre-training objectives either ignore such structure or focus on learning to use it. The feasibility and benefits of learning the knowledge of analyzing semantic structure have not been investigated. To this end, this work proposes two novel pre-training objectives, namely Control Dependency Prediction (CDP) and Data Dependency Prediction (DDP), which aim to predict the statement-level control dependencies and token-level data dependencies, respectively, in a code snippet only based on its source code. During pre-training, CDP and DDP can guide the model to learn the knowledge required for analyzing fine-grained dependencies in code. After pre-training, the pre-trained model can boost the understanding of vulnerable code during fine-tuning and can directly be used to perform dependence analysis for both partial and complete functions. To demonstrate the benefits of our pre-training objectives, we pre-train a Transformer model named PDBERT with CDP and DDP, fine-tune it on three vulnerability analysis tasks, i.e., vulnerability detection, vulnerability classification, and vulnerability assessment, and also evaluate it on program dependence analysis. Experimental results show that PDBERT benefits from CDP and DDP, leading to state-of-the-art performance on the three downstream tasks. Also, PDBERT achieves F1-scores of over 99% and 94% for predicting control and data dependencies, respectively, in partial and complete functions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller 等NDSS 2026 · 被引用 17 次
- Top Score on the Wrong Exam: On Benchmarking in Machine Learning for Vulnerability DetectionNiklas Risse, Jing Liu, Marcel BöhmeISSTA 2025 · 被引用 8 次
- SymRadar: PoC-Centered Bounded Verification for Vulnerability RepairSeungheon Han, YoungJae Kim, Yeseung Lee, Jooyong YiICSE 2026 · 被引用 1 次
- Towards More Trustworthy Deep Code Models by Enabling Out-of-Distribution DetectionYanfu Yan, Viet Duong, Huajie Shao, Denys PoshyvanykICSE 2025 · 被引用 1 次
- Large Language Model-Aided Partial Program Dependence AnalysisXiaokai Rong, Aashish Yadavally, Tien N. NguyenICSE 2026 · 被引用 1 次
它引用的顶会 Paper26
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 被引用 438 次
相关 Paper
- Coding-PTMs: How to Find Optimal Code Pre-trained Models for Code Embedding in Vulnerability Detection?Yu Zhao, Lina Gong, Zhiqiu Huang, Yongwei Wang 等ASE 2024 · 被引用 10 次
- CodeArt: Better Code Models by Attention Regularization When Symbols Are LackingZian Su, Xiangzhe Xu, Ziyang Huang, Zhuo Zhang 等FSE 2024 · 被引用 1 次
- Learning to Locate and Describe VulnerabilitiesJian Zhang, Shangqing Liu, Xu Wang, Tianlin Li 等ASE 2023 · 被引用 8 次
- SCALE: Constructing Structured Natural Language Comment Trees for Software Vulnerability DetectionXin-Cheng Wen, Cuiyun Gao, Shuzheng Gao, Yang Xiao 等ISSTA 2024 · 被引用 17 次
- Towards Learning (Dis)-Similarity of Source Code from Program ContrastsYangruibo Ding, Luca Buratti, Saurabh Pujar, Alessandro Morari 等ACL 2022
