Learning to Complete Code with Sketches
Daya Guo, Alexey Svyatkovskiy, Jian Yin, Nan Duan, Marc Brockschmidt, Miltiadis Allamanis
摘要
Code completion is usually cast as a language modelling problem, i.e., continuing an input in a left-to-right fashion. However, in practice, some parts of the completion (e.g., string literals) may be very hard to predict, whereas subsequent parts directly follow from the context. To handle this, we instead consider the scenario of generating code completions with "holes" inserted in places where a model is uncertain. We develop GRAMMFORMER, a Transformer-based model that guides code generation by the programming language grammar, and compare it to a variety of more standard sequence models. We train the models on code completion for C# and Python given partial code context. To evaluate models, we consider both ROUGE as well as a new metric REGEXACC that measures success of generating completions matching long outputs with as few holes as possible. In our experiments, GRAMMFORMER generates 10-50% more accurate completions compared to traditional generative models and 37-50% longer sketches compared to sketch-generating baselines trained with similar techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese 等NeurIPS 2022 · 被引用 571 次
- Grounded Copilot: How Programmers Interact with Code-Generating ModelsShraddha Barke, Michael B. James, Nadia PolikarpovaOOPSLA 2023 · 被引用 408 次
- ReACC: A Retrieval-Augmented Code Completion FrameworkShuai Lu, Nan Duan, Hojae Han, Daya Guo 等ACL 2022 · 被引用 208 次
- Repair Is Nearly Generation: Multilingual Program Repair with LLMsHarshit Joshi, José Pablo Cambronero Sánchez, Sumit Gulwani, Vu Le 等AAAI 2023 · 被引用 182 次
- Large Language Models of Code Fail at Completing Code with Potential BugsTuan Dinh, Jinman Zhao, Samson Tan, Renato Negrinho 等NeurIPS 2023 · 被引用 59 次
它引用的顶会 Paper6
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis 等ICLR 2020 · 被引用 252 次
- A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text GenerationTianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao 等ACL 2022 · 被引用 194 次
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 被引用 179 次
- Big code != big vocabulary: open-vocabulary models for source codeRafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton 等ICSE 2020 · 被引用 140 次
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 被引用 54 次
相关 Paper
- Language Models for Code Completion: A Practical EvaluationMaliheh Izadi, Jonathan Katzy, Tim van Dam, Marc Otten 等ICSE 2024 · 被引用 51 次
- CodeFill: Multi-token Code Completion by Jointly learning from Structure and Naming SequencesMaliheh Izadi, Roberta Gismondi, Georgios GousiosICSE 2022 · 被引用 79 次
- Empirical study of transformers for source codeNadezhda Chirkova, Sergey TroshinFSE 2021 · 被引用 53 次
- Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax HierarchyColin B. Clement, Shuai Lu, Xiaoyu Liu, Michele Tufano 等EMNLP 2021 · 被引用 11 次
- LongCoder: A Long-Range Pre-trained Language Model for Code CompletionDaya Guo, Canwen Xu, Nan Duan, Jian Yin 等ICML 2023 · 被引用 150 次
