Modeling Hierarchical Syntax Structure with Triplet Position for Source Code Summarization
Juncai Guo, Jin Liu, Yao Wan, Li Li, Pingyi Zhou
摘要
Automatic code summarization, which aims to describe the source code in natural language, has become an essential task in software maintenance. Our fellow researchers have attempted to achieve such a purpose through various machine learning-based approaches. One key challenge keeping these approaches from being practical lies in the lacking of retaining the semantic structure of source code, which has unfortunately been overlooked by the stateof-the-art methods. Existing approaches resort to representing the syntax structure of code by modeling the Abstract Syntax Trees (ASTs). However, the hierarchical structures of ASTs have not been well explored. In this paper, we propose CODESCRIBE to model the hierarchical syntax structure of code by introducing a novel triplet position for code summarization. Specifically, CODESCRIBE leverages the graph neural network and Transformer to preserve the structural and sequential information of code, respectively. In addition, we propose a pointer-generator network that pays attention to both the structure and sequential tokens of code for a better summary generation. Experiments on two real-world datasets in Java and Python demonstrate the effectiveness of our proposed approach when compared with several state-of-the-art baselines. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo CodeTong Ye, Lingfei Wu, Tengfei Ma, Xuhong Zhang 等EMNLP 2023 · 被引用 4 次
- ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation GroundingIndraneil Paul, Haoyi Yang, Goran Glavas, Kristian Kersting 等ICLR 2025
它引用的顶会 Paper4
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Leveraging Code Generation to Improve Code Retrieval and Summarization via Dual LearningWei Ye, Rui Xie, Jinglei Zhang, Tianxiang Hu 等WWW 2020 · 被引用 83 次
- Retrieve and Refine: Exemplar-based Neural Comment GenerationBolin Wei, Yongmin Li, Ge Li, Xin Xia 等ASE 2020 · 被引用 68 次
相关 Paper
- CAST: Enhancing Code Summarization with Hierarchical Splitting and Reconstruction of Abstract Syntax TreesEnsheng Shi, Yanlin Wang, Lun Du, Hongyu Zhang 等EMNLP 2021 · 被引用 42 次
- AST-Trans: Code Summarization with Efficient Tree-Structured AttentionZe Tang, Xiaoyu Shen, Chuanyi Li, Jidong Ge 等ICSE 2022 · 被引用 57 次
- MGF-ESE: An Enhanced Semantic Extractor with Multi-Granularity Feature Fusion for Code SummarizationXiaolong Xu, Yuxin Cao, Hongsheng Hu, Haolong Xiang 等WWW 2025 · 被引用 4 次
- Retrieval-Augmented Generation for Code Summarization via Hybrid GNNShangqing Liu, Yu Chen, Xiaofei Xie, Jing Kai Siow 等ICLR 2021 · 被引用 194 次
- Retrieval-based neural source code summarizationJian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun 等ICSE 2020 · 被引用 242 次
