GrammarT5: Grammar-Integrated Pretrained Encoder-Decoder Neural Model for Code
Qihao Zhu, Qingyuan Liang, Zeyu Sun, Yingfei Xiong, Lu Zhang, Shengyu Cheng
摘要
Pretrained models for code have exhibited promising performance across various code-related tasks, such as code summarization, code completion, code translation, and bug detection. However, despite their success, the majority of current models still represent code as a token sequence, which may not adequately capture the essence of the underlying code structure. In this work, we propose GrammarT5, a grammar-integrated encoder-decoder pretrained neural model for code. GrammarT5 employs a novel grammar-integrated representation, Tokenized Grammar Rule Sequence (TGRS), for code. TGRS is constructed based on the grammar rule sequence utilized in syntax-guided code generation and integrates syntax information with code tokens within an appropriate input length. Furthermore, we suggest attaching language flags to help GrammarT5 differentiate between grammar rules of various programming languages. Finally, we introduce two novel pretraining tasks-Edge Prediction (EP), and Sub-Tree Prediction (STP) to learn syntactic information. Experiments were conducted on five code-related tasks using eleven datasets, demonstrating that GrammarT5 achieves stateof-the-art (SOTA) performance on most tasks in comparison to models of the same scale. Additionally, the paper illustrates that the proposed pretraining tasks and language flags can enhance GrammarT5 to better capture the syntax and semantics of code.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Towards Understanding the Characteristics of Code Generation Errors Made by Large Language ModelsZhijie Wang, Zijie Zhou, Da Song, Yuheng Huang 等ICSE 2025 · 被引用 12 次
- Mutual Learning-Based Framework for Enhancing Robustness of Code Models via Adversarial TrainingYangsen Wang, Yizhou Chen, Yifan Zhao, Zhihao Gong 等ASE 2024 · 被引用 3 次
- Speculative Decoding for Verilog: Speed and Quality, All in OneChangran Xu, Yi Liu, Yunhao Zhou, Shan Huang 等DAC 2025 · 被引用 1 次
- Chiseling Out Efficiency: Structured Skeleton Supervision for Efficient Code GenerationYu Yu, Zhihong Sun, Jia Li, Yao Wan 等FSE 2026
- Contamination Means Overestimation? A Fine-Grained Empirical Study in Code IntelligenceZhen Yang, Hongyi Lin, Yifan He, Junqi Wang 等ISSTA 2026
它引用的顶会 Paper15
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese 等NeurIPS 2022 · 被引用 571 次
相关 Paper
- SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code RepresentationsChangan Niu, Chuanyi Li, Vincent Ng, Jidong Ge 等ICSE 2022 · 被引用 99 次
- AST-T5: Structure-Aware Pretraining for Code Generation and UnderstandingLinyuan Gong, Mostafa Elhoushi, Alvin CheungICML 2024 · 被引用 42 次
- What Do They Capture? - A Structural Analysis of Pre-Trained Language Models for Source CodeYao Wan, Wei Zhao, Hongyu Zhang, Yulei Sui 等ICSE 2022 · 被引用 66 次
- CoditT5: Pretraining for Source Code and Natural Language EditingJiyang Zhang, Sheena Panthaplackel, Pengyu Nie, Junyi Jessy Li 等ASE 2022 · 被引用 81 次
- FAIR: Flow Type-Aware Pre-Training of Compiler Intermediate RepresentationsChangan Niu, Chuanyi Li, Vincent Ng, David Lo 等ICSE 2024 · 被引用 5 次
