CodeFill: Multi-token Code Completion by Jointly learning from Structure and Naming Sequences
Maliheh Izadi, Roberta Gismondi, Georgios Gousios
Abstract
Code completion is an essential feature of IDEs, yet current autocompleters are restricted to either grammar-based or NLP-based single token completions. Both approaches have significant drawbacks: grammar-based autocompletion is restricted in dynamicallytyped language environments, whereas NLP-based autocompleters struggle to understand the semantics of the programming language and the developer's code context. In this work, we present CodeFill, a language model for autocompletion that combines learned structure and naming information. Using a parallel Transformer architecture and multi-task learning, CodeFill consumes sequences of source code token names and their equivalent AST token types. Uniquely, CodeFill is trained both for single-token and multi-token (statement) prediction, which enables it to learn long-range dependencies among grammatical and naming elements. We train CodeFill on two datasets, consisting of 29M and 425M lines of code, respectively. To make the evaluation more realistic, we develop a method to automatically infer points in the source code at which completion matters. We compare CodeFill against four baselines and two state-of-the-art models, GPT-C and TravTrans+. CodeFill surpasses all baselines in single token prediction (MRR: 70.9% vs. 66.2% and 67.8%) and outperforms the state of the art for multi-token prediction (ROUGE-L: 63.7% vs. 52.4% and 59.2%, for 𝑛 = 4 tokens). We publicly release our source code and datasets. CCS CONCEPTS • Software and its engineering → Software notations and tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- InCoder: A Generative Model for Code Infilling and SynthesisDaniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang et al.ICLR 2023 · 140 citations
- Language Models for Code Completion: A Practical EvaluationMaliheh Izadi, Jonathan Katzy, Tim van Dam, Marc Otten et al.ICSE 2024 · 51 citations
- DeepVD: Toward Class-Separation Features for Neural Network Vulnerability DetectionWenbo Wang, Tien N. Nguyen, Shaohua Wang, Yi Li et al.ICSE 2023 · 32 citations
- Code Search is All You Need? Improving Code Suggestions with Code SearchJunkai Chen, Xing Hu, Zhenhao Li, Cuiyun Gao et al.ICSE 2024 · 31 citations
- "Len or index or count, anything but v1": Predicting Variable Names in Decompilation Output with Transfer LearningKuntal Kumar Pal, Ati Priya Bajaj, Pratyay Banerjee, Audrey Dutcher et al.S&P 2024 · 30 citations
Builds on5
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 179 citations
- End-to-End Neural Pipeline for Goal-Oriented Dialogue Systems using GPT-2DongHoon Ham, Jeong-Gwan Lee, Youngsoo Jang, Kee-Eung KimACL 2020 · 167 citations
- Multi-task Learning based Pre-trained Language Model for Code CompletionFang Liu, Ge Li, Yunfei Zhao, Zhi JinASE 2020 · 162 citations
- Big code != big vocabulary: open-vocabulary models for source codeRafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton et al.ICSE 2020 · 140 citations
- Siri, Write the Next MethodFengcai Wen, Emad Aghajani, Csaba Nagy, Michele Lanza et al.ICSE 2021
Related papers
- Learning to Complete Code with SketchesDaya Guo, Alexey Svyatkovskiy, Jian Yin, Nan Duan et al.ICLR 2022 · 42 citations
- ReACC: A Retrieval-Augmented Code Completion FrameworkShuai Lu, Nan Duan, Hojae Han, Daya Guo et al.ACL 2022 · 208 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- McEval: Massively Multilingual Code EvaluationLinzheng Chai, Shukai Liu, Jian Yang, Yuwei Yin et al.ICLR 2025 · 1 citation
