When Less is Enough: Positive and Unlabeled Learning Model for Vulnerability Detection
Xin-Cheng Wen, Xinchen Wang, Cuiyun Gao, Shaohua Wang, Yang Liu, Zhaoquan Gu
摘要
Automated code vulnerability detection has gained increasing attention in recent years. The deep learning (DL)-based methods, which implicitly learn vulnerable code patterns, have proven effective in vulnerability detection. The performance of DL-based methods usually relies on the quantity and quality of labeled data. However, the current labeled data are generally automatically collected, such as crawled from human-generated commits, making it hard to ensure the quality of the labels. Prior studies have demonstrated that the non-vulnerable code (i.e., negative labels) tends to be unreliable in commonly-used datasets, while vulnerable code (i.e., positive labels) is more determined. Considering the large numbers of unlabeled data in practice, it is necessary and worth exploring to leverage the positive data and large numbers of unlabeled data for more accurate vulnerability detection. In this paper, we focus on the Positive and Unlabeled (PU) learning problem for vulnerability detection and propose a novel model named PILOT, i.e., Positive and unlabeled Learning mOdel for vulnerability deTection. PILOT only learns from positive and unlabeled data for vulnerability detection. It mainly contains two modules: (1) A distance-aware label selection module, aiming at generating pseudo-labels for selected unlabeled data, which involves the inter-class distance prototype and progressive fine-tuning; (2) A mixed-supervision representation learning module to further alleviate the influence of noise and enhance the discrimination of representations. Extensive experiments in vulnerability detection are conducted to evaluate the effectiveness of PILOT based on real-world vulnerability datasets. The experimental results show that PILOT outperforms the popular weakly supervised methods by 2.78%-18.93% in the PU learning setting. Compared with the state-of-the-art methods, PILOT also improves the performance of 1.34%-12.46 % in F1 score metrics in the supervised setting. In addition, PILOT can identify 23 mislabeled from the FFMPeg+Qemu dataset in the PU learning setting based on manual checking.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- SCALE: Constructing Structured Natural Language Comment Trees for Software Vulnerability DetectionXin-Cheng Wen, Cuiyun Gao, Shuzheng Gao, Yang Xiao 等ISSTA 2024 · 被引用 17 次
- SelfPiCo: Self-Guided Partial Code Execution with LLMsZhipeng Xue, Zhipeng Gao, Shaohua Wang, Xing Hu 等ISSTA 2024 · 被引用 8 次
- Your Fix Is My Exploit: Enabling Comprehensive DL Library API Fuzzing with Large Language ModelsKunpeng Zhang, Shuai Wang, Jitao Han, Xiaogang Zhu 等ICSE 2025 · 被引用 6 次
- Bridge and Hint: Extending Pre-trained Language Models for Long-Range CodeYujia Chen, Cuiyun Gao, Zezhou Yang, Hongyu Zhang 等ISSTA 2024 · 被引用 4 次
- Repository-Level Graph Representation Learning for Enhanced Security Patch DetectionXin-Cheng Wen, Zirui Lin, Cuiyun Gao, Hongyu Zhang 等ICSE 2025 · 被引用 1 次
它引用的顶会 Paper16
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 被引用 283 次
- No more fine-tuning? an experimental evaluation of prompt tuning in code intelligenceChaozheng Wang, Yuanhang Yang, Cuiyun Gao, Yun Peng 等FSE 2022 · 被引用 148 次
- VulCNN: An Image-inspired Scalable Vulnerability Detection SystemYueming Wu, Deqing Zou, Shihan Dou, Wei Yang 等ICSE 2022 · 被引用 141 次
- ProGraML: A Graph-based Program Representation for Data Flow Analysis and Compiler OptimizationsChris Cummins, Zacharias V. Fisches, Tal Ben-Nun, Torsten Hoefler 等ICML 2021 · 被引用 140 次
- Data Quality for Software Vulnerability DatasetsRoland Croft, Muhammad Ali Babar, M. Mehdi KholoosiICSE 2023 · 被引用 138 次
相关 Paper
- VulSim: Leveraging Similarity of Multi-Dimensional Neighbor Embeddings for Vulnerability DetectionSamiha Shimmi, Ashiqur Rahman, Mohan Gadde, Hamed Okhravi 等USENIX Security 2024 · 被引用 13 次
- An Empirical Study on Noisy Label Learning for Program UnderstandingWenhan Wang, Yanzhou Li, Anran Li, Jian Zhang 等ICSE 2024 · 被引用 5 次
- Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code ModelsShuzheng Gao, Wenxin Mao, Cuiyun Gao, Li Li 等ICSE 2024 · 被引用 15 次
- An Empirical Study of Deep Learning Models for Vulnerability DetectionBenjamin Steenhoek, Md Mahbubur Rahman, Richard Jiles, Wei LeICSE 2023 · 被引用 107 次
- GVI: Guided Vulnerability Imagination for Boosting Deep Vulnerability DetectorsHeng Yong, Zhong Li, Minxue Pan, Tian Zhang 等ICSE 2025 · 被引用 2 次
