Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax Hierarchy
Colin B. Clement, Shuai Lu, Xiaoyu Liu, Michele Tufano, Dawn Drain, Nan Duan, Neel Sundaresan, Alexey Svyatkovskiy
Abstract
Statistical language modeling and translation with transformers have found many successful applications in program understanding and generation tasks, setting high benchmarks for tools in modern software development environments. The finite context window of these neural models means, however, that they will be unable to leverage the entire relevant context of large files and packages for any given task. While there are many efforts to extend the context window, we introduce an architectureindependent approach for leveraging the syntactic hierarchies of source code for incorporating entire file-level context into a fixedlength window. Using concrete syntax trees of each source file we extract syntactic hierarchies and integrate them into context window by selectively removing from view more specific, less relevant scopes for a given task. We evaluate this approach on code generation tasks and joint translation of natural language and source code in Python programming language, achieving a new state-of-the-art in code completion and summarization for Python in the CodeXGLUE benchmark. We also introduce new CodeXGLUE benchmarks for userexperience-motivated tasks: code completion with normalized literals, method body completion/code summarization conditioned on filelevel context.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a5a17ec-90f7-43b3-a9ee-59393c05d80eCited by top-tier papers8
- LongCoder: A Long-Range Pre-trained Language Model for Code CompletionDaya Guo, Canwen Xu, Nan Duan, Jian Yin et al.ICML 2023 · 150 citations
- Better Context Makes Better Code Language Models: A Case Study on Function Call Argument CompletionHengzhi Pei, Jinman Zhao, Leonard Lausen, Sheng Zha et al.AAAI 2023 · 30 citations
- Multi-lingual Evaluation of Code Generation ModelsBen Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang, Xiaopeng Li et al.ICLR 2023 · 28 citations
- CodeTrek: Flexible Modeling of Code using an Extensible Relational RepresentationPardis Pashakhanloo, Aaditya Naik, Yuepeng Wang, Hanjun Dai et al.ICLR 2022 · 15 citations
- Bridge and Hint: Extending Pre-trained Language Models for Long-Range CodeYujia Chen, Cuiyun Gao, Zezhou Yang, Hongyu Zhang et al.ISSTA 2024 · 4 citations
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Pre-training via ParaphrasingMike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan et al.NeurIPS 2020 · 165 citations
- CPC: automatically classifying and propagating natural language comments via program analysisJuan Zhai, Xiangzhe Xu, Yu Shi, Guanhong Tao et al.ICSE 2020 · 39 citations
Related papers
- ReACC: A Retrieval-Augmented Code Completion FrameworkShuai Lu, Nan Duan, Hojae Han, Daya Guo et al.ACL 2022 · 208 citations
- Empirical study of transformers for source codeNadezhda Chirkova, Sergey TroshinFSE 2021 · 53 citations
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 438 citations
- Modeling Hierarchical Syntax Structure with Triplet Position for Source Code SummarizationJuncai Guo, Jin Liu, Yao Wan, Li Li et al.ACL 2022
- PyMT5: multi-mode translation of natural language and Python code with transformersColin B. Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy et al.EMNLP 2020 · 24 citations
