AST-T5: Structure-Aware Pretraining for Code Generation and Understanding
Linyuan Gong, Mostafa Elhoushi, Alvin Cheung
摘要
Large language models (LLMs) have made significant advancements in code-related tasks, yet many LLMs treat code as simple sequences, neglecting its structured nature. We introduce AST-T5, a novel pretraining paradigm that leverages the Abstract Syntax Tree (AST) for enhanced code generation, transpilation, and understanding. Using dynamic programming, our AST-Aware Segmentation retains code structure, while our AST-Aware Span Corruption objective equips the model to reconstruct various code structures. Unlike other models, AST-T5 avoids complex program analyses or architectural changes, so it integrates seamlessly with any encoderdecoder Transformer. Evaluations show that AST-T5 consistently outperforms similar-sized LMs across various code-related tasks including HumanEval and MBPP. Structure-awareness makes AST-T5 particularly powerful in code-tocode tasks, surpassing CodeT5 by 2 points in exact match score for the Bugs2Fix task and by 3 points in exact match score for Java-C# Transpilation in CodeXGLUE. Our code and model are publicly available at https://github.com/ gonglinyuan/ast t5 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle TasksLinyuan Gong, Sida Wang, Mostafa Elhoushi, Alvin CheungICML 2024 · 被引用 33 次
- ExeCoder: Empowering Large Language Models with Executability Representation for Code TranslationMinghua He, Yue Chen, Fangkai Yang, Pu Zhao 等EMNLP 2025 · 被引用 1 次
- ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation GroundingIndraneil Paul, Haoyi Yang, Goran Glavas, Kristian Kersting 等ICLR 2025
- Alignment with Fill-In-the-Middle for Enhancing Code GenerationHouxing Ren, Zimu Lu, Weikang Shi, Haotian Hou 等EMNLP 2025
- ARBench: Algorithmic Reasoner or API Alchemist? Evaluating LLMs Beyond API CallsRenbiao Liu, Chao-Zeng Ma, Anqi Li, Hui Sun 等AAAI 2026
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
相关 Paper
- SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code RepresentationsChangan Niu, Chuanyi Li, Vincent Ng, Jidong Ge 等ICSE 2022 · 被引用 99 次
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui 等EMNLP 2023 · 被引用 339 次
- GrammarT5: Grammar-Integrated Pretrained Encoder-Decoder Neural Model for CodeQihao Zhu, Qingyuan Liang, Zeyu Sun, Yingfei Xiong 等ICSE 2024 · 被引用 10 次
- AST-Probe: Recovering abstract syntax trees from hidden representations of pre-trained language modelsJosé Antonio Hernández López, Martin Weyssow, Jesús Sánchez Cuadrado, Houari A. SahraouiASE 2022 · 被引用 18 次
- CoditT5: Pretraining for Source Code and Natural Language EditingJiyang Zhang, Sheena Panthaplackel, Pengyu Nie, Junyi Jessy Li 等ASE 2022 · 被引用 81 次
