Code Prediction by Feeding Trees to Transformers
Seohyun Kim, Jinman Zhao, Yuchi Tian, Satish Chandra
Abstract
Code prediction, more specifically autocomplete, has become an essential feature in modern IDEs. Autocomplete is more effective when the desired next token is at (or close to) the top of the list of potential completions offered by the IDE at cursor position. This is where the strength of the underlying machine learning system that produces a ranked order of potential completions comes into play. We advance the state-of-the-art in the accuracy of code prediction (next token prediction) used in autocomplete systems. Our work uses Transformers as the base neural architecture. We show that by making the Transformer architecture aware of the syntactic structure of code, we increase the margin by which a Transformer-based system outperforms previous systems. With this, it outperforms the accuracy of several state-of-the-art next token prediction systems by margins ranging from 14% to 18%. We present in the paper several ways of communicating the code structure to the Transformer, which is fundamentally built for processing sequence data. We provide a comprehensive experimental evaluation of our proposal, along with alternative design choices, on a standard Python dataset, as well as on Facebook internal Python corpus. Our code and data preparation pipeline will be available in open source.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc977762-8a9e-402f-967f-15da79293116Cited by top-tier papers54
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- ReACC: A Retrieval-Augmented Code Completion FrameworkShuai Lu, Nan Duan, Hojae Han, Daya Guo et al.ACL 2022 · 208 citations
- DOBF: A Deobfuscation Pre-Training Objective for Programming LanguagesMarie-Anne Lachaux, Baptiste Rozière, Marc Szafraniec, Guillaume LampleNeurIPS 2021 · 174 citations
- Multi-task Learning based Pre-trained Language Model for Code CompletionFang Liu, Ge Li, Yunfei Zhao, Zhi JinASE 2020 · 162 citations
- Leveraging Automated Unit Tests for Unsupervised Code TranslationBaptiste Rozière, Jie Zhang, François Charton, Mark Harman et al.ICLR 2022 · 161 citations
Builds on5
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis et al.ICLR 2020 · 252 citations
- Adversarial examples for models of codeNoam Yefet, Uri Alon, Eran YahavOOPSLA 2020 · 162 citations
- Big code != big vocabulary: open-vocabulary models for source codeRafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton et al.ICSE 2020 · 140 citations
- Structural Language Models of CodeUri Alon, Roy Sadaka, Omer Levy, Eran YahavICML 2020 · 115 citations
- Tree-Structured Attention with Hierarchical AccumulationXuan-Phi Nguyen, Shafiq R. Joty, Steven C. H. Hoi, Richard SocherICLR 2020 · 79 citations
Related papers
- Language Models for Code Completion: A Practical EvaluationMaliheh Izadi, Jonathan Katzy, Tim van Dam, Marc Otten et al.ICSE 2024 · 51 citations
- Empirical study of transformers for source codeNadezhda Chirkova, Sergey TroshinFSE 2021 · 53 citations
- A structural model for contextual code changesShaked Brody, Uri Alon, Eran YahavOOPSLA 2020 · 36 citations
- CodeFill: Multi-token Code Completion by Jointly learning from Structure and Naming SequencesMaliheh Izadi, Roberta Gismondi, Georgios GousiosICSE 2022 · 79 citations
- Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax HierarchyColin B. Clement, Shuai Lu, Xiaoyu Liu, Michele Tufano et al.EMNLP 2021 · 11 citations
