Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax Hierarchy
Colin B. Clement, Shuai Lu, Xiaoyu Liu, Michele Tufano, Dawn Drain, Nan Duan, Neel Sundaresan, Alexey Svyatkovskiy
摘要
Statistical language modeling and translation with transformers have found many successful applications in program understanding and generation tasks, setting high benchmarks for tools in modern software development environments. The finite context window of these neural models means, however, that they will be unable to leverage the entire relevant context of large files and packages for any given task. While there are many efforts to extend the context window, we introduce an architectureindependent approach for leveraging the syntactic hierarchies of source code for incorporating entire file-level context into a fixedlength window. Using concrete syntax trees of each source file we extract syntactic hierarchies and integrate them into context window by selectively removing from view more specific, less relevant scopes for a given task. We evaluate this approach on code generation tasks and joint translation of natural language and source code in Python programming language, achieving a new state-of-the-art in code completion and summarization for Python in the CodeXGLUE benchmark. We also introduce new CodeXGLUE benchmarks for userexperience-motivated tasks: code completion with normalized literals, method body completion/code summarization conditioned on filelevel context.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- LongCoder: A Long-Range Pre-trained Language Model for Code CompletionDaya Guo, Canwen Xu, Nan Duan, Jian Yin 等ICML 2023 · 被引用 150 次
- Better Context Makes Better Code Language Models: A Case Study on Function Call Argument CompletionHengzhi Pei, Jinman Zhao, Leonard Lausen, Sheng Zha 等AAAI 2023 · 被引用 30 次
- Multi-lingual Evaluation of Code Generation ModelsBen Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang, Xiaopeng Li 等ICLR 2023 · 被引用 28 次
- CodeTrek: Flexible Modeling of Code using an Extensible Relational RepresentationPardis Pashakhanloo, Aaditya Naik, Yuepeng Wang, Hanjun Dai 等ICLR 2022 · 被引用 15 次
- Bridge and Hint: Extending Pre-trained Language Models for Long-Range CodeYujia Chen, Cuiyun Gao, Zezhou Yang, Hongyu Zhang 等ISSTA 2024 · 被引用 4 次
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Pre-training via ParaphrasingMike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan 等NeurIPS 2020 · 被引用 165 次
- CPC: automatically classifying and propagating natural language comments via program analysisJuan Zhai, Xiangzhe Xu, Yu Shi, Guanhong Tao 等ICSE 2020 · 被引用 39 次
相关 Paper
- ReACC: A Retrieval-Augmented Code Completion FrameworkShuai Lu, Nan Duan, Hojae Han, Daya Guo 等ACL 2022 · 被引用 208 次
- Empirical study of transformers for source codeNadezhda Chirkova, Sergey TroshinFSE 2021 · 被引用 53 次
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 被引用 438 次
- Modeling Hierarchical Syntax Structure with Triplet Position for Source Code SummarizationJuncai Guo, Jin Liu, Yao Wan, Li Li 等ACL 2022
- PyMT5: multi-mode translation of natural language and Python code with transformersColin B. Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy 等EMNLP 2020 · 被引用 24 次
