AST-Trans: Code Summarization with Efficient Tree-Structured Attention
Ze Tang, Xiaoyu Shen, Chuanyi Li, Jidong Ge, Liguo Huang, Zheling Zhu, Bin Luo
Abstract
Code summarization aims to generate brief natural language descriptions for source codes. The state-of-the-art approaches follow a transformer-based encoder-decoder architecture. As the source code is highly structured and follows strict grammars, its Abstract Syntax Tree (AST) is widely used for encoding structural information. However, ASTs are much longer than the corresponding source code. Existing approaches ignore the size constraint and simply feed the whole linearized AST into the encoders. We argue that such a simple process makes it difficult to extract the truly useful dependency relations from the overlong input sequence. It also incurs significant computational overhead since each node needs to apply self-attention to all other nodes in the AST. To encode the AST more effectively and efficiently, we propose AST-Trans in this paper which exploits two types of node relationships in the AST: ancestor-descendant and sibling relationships. It applies the tree-structured attention to dynamically allocate weights for relevant nodes and exclude irrelevant nodes based on these two relationships. We further propose an efficient implementation to support fast parallel computation for tree-structure attention. On the two code summarization datasets, experimental results show that AST-Trans significantly outperforms the state-of-the-arts while being times more efficient than standard transformers 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d048a99a-1a1e-4ef0-bea6-007b1e16797fCited by top-tier papers9
- Tare: Type-Aware Neural Program RepairQihao Zhu, Zeyu Sun, Wenjie Zhang, Yingfei Xiong et al.ICSE 2023 · 29 citations
- Developer-Intent Driven Code Comment GenerationFangwen Mu, Xiao Chen, Lin Shi, Song Wang et al.ICSE 2023 · 25 citations
- Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code ModelsShuzheng Gao, Wenxin Mao, Cuiyun Gao, Li Li et al.ICSE 2024 · 15 citations
- EyeTrans: Merging Human and Machine Attention for Neural Code SummarizationYifan Zhang, Jiliang Li, Zachary Karas, Aakash Bansal et al.FSE 2024 · 15 citations
- Keeping Pace with Ever-Increasing Data: Towards Continual Learning of Code Intelligence ModelsShuzheng Gao, Hongyu Zhang, Cuiyun Gao, Chaozheng WangICSE 2023 · 14 citations
Builds on5
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 179 citations
- Language-Agnostic Representation Learning of Source Code from Structure and ContextDaniel Zügner, Tobias Kirschstein, Michele Catasta, Jure Leskovec et al.ICLR 2021 · 131 citations
- MovieChats: Chat like Humans in a Closed DomainHui Su, Xiaoyu Shen, Xiao Zhou, Zheng Zhang et al.EMNLP 2020 · 24 citations
Related papers
- CAST: Enhancing Code Summarization with Hierarchical Splitting and Reconstruction of Abstract Syntax TreesEnsheng Shi, Yanlin Wang, Lun Du, Hongyu Zhang et al.EMNLP 2021 · 42 citations
- Integrating Tree Path in Transformer for Code RepresentationHan Peng, Ge Li, Wenhan Wang, Yunfei Zhao et al.NeurIPS 2021 · 56 citations
- MGF-ESE: An Enhanced Semantic Extractor with Multi-Granularity Feature Fusion for Code SummarizationXiaolong Xu, Yuxin Cao, Hongsheng Hu, Haolong Xiang et al.WWW 2025 · 4 citations
- Modeling Hierarchical Syntax Structure with Triplet Position for Source Code SummarizationJuncai Guo, Jin Liu, Yao Wan, Li Li et al.ACL 2022
- Rethinking Positional Encoding in Tree Transformer for Code RepresentationHan Peng, Ge Li, Yunfei Zhao, Zhi JinEMNLP 2022 · 10 citations
