An Empirical Comparison of Pre-Trained Models of Source Code
Changan Niu, Chuanyi Li, Vincent Ng, Dongxiao Chen, Jidong Ge, Bin Luo
Abstract
While a large number of pre-trained models of source code have been successfully developed and applied to a variety of software engineering (SE) tasks in recent years, our understanding of these pre-trained models is arguably fairly limited. With the goal of advancing our understanding of these models, we perform the first systematic empirical comparison of 19 recently-developed pre-trained models of source code on 13 SE tasks. To gain additional insights into these models, we adopt a recently-developed 4-dimensional categorization of pretrained models, and subsequently investigate whether there are correlations between different categories of pre-trained models and their performances on different SE tasks. Index Terms-Pre-training of Source Code, AI for SE TABLE I DETAILS OF EVALUATION TASKS, DATASETS AND METRICS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- An Empirical Study on Fine-Tuning Large Language Models of Code for Automated Program RepairKai Huang, Xiangxin Meng, Jian Zhang, Yang Liu et al.ASE 2023 · 91 citations
- On the Evaluation of Large Language Models in Unit Test GenerationLin Yang, Chen Yang, Shutao Gao, Weijing Wang et al.ASE 2024 · 42 citations
- Unveiling Memorization in Code ModelsZhou Yang, Zhipeng Zhao, Chenyu Wang, Jieke Shi et al.ICSE 2024 · 33 citations
- Out of Sight, Out of Mind: Better Automatic Vulnerability Repair by Broadening Input Ranges and SourcesXin Zhou, Kisub Kim, Bowen Xu, DongGyun Han et al.ICSE 2024 · 32 citations
- CodeGen4Libs: A Two-Stage Approach for Library-Oriented Code GenerationMingwei Liu, Tianyong Yang, Yiling Lou, Xueying Du et al.ASE 2023 · 31 citations
Builds on21
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 606 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
Related papers
- A Large-Scale Study of Model Integration in ML-Enabled Software SystemsYorick Sens, Henriette Knopp, Sven Peldszus, Thorsten BergerICSE 2025 · 3 citations
- Natural Language to Code: How Far Are We?Shangwen Wang, Mingyang Geng, Bo Lin, Zhensu Sun et al.FSE 2023 · 24 citations
- On Calibration of Pre-trained Code ModelsZhenhao Zhou, Chaofeng Sha, Xin PengICSE 2024 · 3 citations
- Automating Code-Related Tasks Through Transformers: The Impact of Pre-trainingRosalia Tufano, Luca Pascarella, Gabriele BavotaICSE 2023 · 15 citations
- Bridging Pre-trained Models and Downstream Tasks for Source Code UnderstandingDeze Wang, Zhouyang Jia, Shanshan Li, Yue Yu et al.ICSE 2022 · 68 citations
