Coding-PTMs: How to Find Optimal Code Pre-trained Models for Code Embedding in Vulnerability Detection?
Yu Zhao, Lina Gong, Zhiqiu Huang, Yongwei Wang, Mingqiang Wei, Fei Wu
Abstract
Vulnerability detection is garnering increasing attention in software engineering, since code vulnerabilities possibly pose significant security. Recently, reusing various code pre-trained models (e.g., CodeBERT, CodeT5, and CodeGen) has become common for code embedding without providing reasonable justifications in vulnerability detection. The premise for casually utilizing pre-trained models (PTMs) is that the code embeddings generated by different PTMs would generate a similar impact on the performance. Is that TRUE? To answer this important question, we systematically investigate the effects of code embedding generated by ten different code PTMs on the performance of vulnerability detection, and get the answer, i.e., that is NOT true. We observe that code embedding generated by various code PTMs can indeed influence the performance and selecting an embedding technique based on parameter scales and embedding dimension is not reliable. Our findings highlight the necessity of quantifying and evaluating the characteristics of code embedding generated by various code PTMs to understand the effects. To achieve this goal, we analyze the numerical representation and data distribution of code embedding generated by different PTMs to evaluate differences and characteristics. Based on these insights, we propose Coding-PTMs, a recommendation framework to assist engineers in selecting optimal code PTMs for their specific vulnerability detection tasks. Specifically, we define thirteen code embedding metrics across three dimensions (i.e., statistics, norm, and distribution) for constructing a specialized code PTM recommendation dataset. We then employ a Random Forest classifier to train a recommendation model and identify the optimal code PTMs from the candidate model zoo. We encourage engineers to use our Coding-PTMs to evaluate the characteristics of code embeddings generated by candidate code PTMs on the performance and recommend optimal code PTMs for code embedding in their vulnerability detection tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30d9774a-ab9d-4ab2-8f32-ce155a4a6d8dCited by top-tier papers2
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller et al.NDSS 2026 · 17 citations
- Three Heads Are Better Than One: A Multi-perspective Reasoning Framework for Enhanced Vulnerability DetectionXin Peng, Bo Lin, Jing Wang, Xiaoling Li et al.FSE 2026 · 1 citation
Builds on13
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 438 citations
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis et al.ICLR 2020 · 252 citations
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu et al.ICLR 2023 · 234 citations
Related papers
- SCALE: Constructing Structured Natural Language Comment Trees for Software Vulnerability DetectionXin-Cheng Wen, Cuiyun Gao, Shuzheng Gao, Yang Xiao et al.ISSTA 2024 · 17 citations
- Pre-training by Predicting Program Dependencies for Vulnerability Analysis TasksZhongxin Liu, Zhijie Tang, Junwei Zhang, Xin Xia et al.ICSE 2024 · 15 citations
- LLM-based Vulnerability Discovery through the Lens of Code MetricsFelix Weissberg, Lukas Pirch, Erik Imgrund, Jonas Möller et al.ICSE 2026
- CTX-Coder: Cross-Attention Architectures Empower LLMs for Long-Context Vulnerability DetectionJujie Wang, Kangfeng Zheng, Bin Wu, Chunhua Wu et al.AAAI 2026
- Compressing Pre-trained Models of Code into 3 MBJieke Shi, Zhou Yang, Bowen Xu, Hong Jin Kang et al.ASE 2022 · 42 citations
