Retrieval-based neural source code summarization
Jian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun, Xudong Liu
摘要
Source code summarization aims to automatically generate concise summaries of source code in natural language texts, in order to help developers better understand and maintain source code. Traditional work generates a source code summary by utilizing information retrieval techniques, which select terms from original source code or adapt summaries of similar code snippets. Recent studies adopt Neural Machine Translation techniques and generate summaries from code snippets using encoder-decoder neural networks. The neural-based approaches prefer the high-frequency words in the corpus and have trouble with the low-frequency ones. In this paper, we propose a retrieval-based neural source code summarization approach where we enhance the neural model with the most similar code snippets retrieved from the training set. Our approach can take advantages of both neural and retrieval-based techniques. Specifically, we first train an attentional encoder-decoder model based on the code snippets and the summaries in the training set; Second, given one input code snippet for testing, we retrieve its two most similar code snippets in the training set from the aspects of syntax and semantics, respectively; Third, we encode the input and two retrieved code snippets, and predict the summary by fusing them during decoding. We conduct extensive experiments to evaluate our approach and the experimental results show that our proposed approach can improve the state-of-the-art methods. CCS CONCEPTS • Software and its engineering → Software maintenance tools.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper50
- Retrieval-Augmented Generation for Code Summarization via Hybrid GNNShangqing Liu, Yu Chen, Xiaofei Xie, Jing Kai Siow 等ICLR 2021 · 被引用 194 次
- An extensive study on pre-trained models for program understanding and generationZhengran Zeng, Hanzhuo Tan, Haotian Zhang, Jing Li 等ISSTA 2022 · 被引用 142 次
- Large Language Models are Few-Shot Summarizers: Multi-Intent Comment Generation via In-Context LearningMingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang 等ICSE 2024 · 被引用 124 次
- Reassessing automatic evaluation metrics for code summarization tasksDevjeet Roy, Sarah Fakhoury, Venera ArnaoudovaFSE 2021 · 被引用 103 次
- SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code RepresentationsChangan Niu, Chuanyi Li, Vincent Ng, Jidong Ge 等ICSE 2022 · 被引用 99 次
相关 Paper
- EditSum: A Retrieve-and-Edit Framework for Source Code SummarizationJia Li, Yongmin Li, Ge Li, Xing Hu 等ASE 2021 · 被引用 55 次
- Retrieve and Refine: Exemplar-based Neural Comment GenerationBolin Wei, Yongmin Li, Ge Li, Xin Xia 等ASE 2020 · 被引用 68 次
- EyeTrans: Merging Human and Machine Attention for Neural Code SummarizationYifan Zhang, Jiliang Li, Zachary Karas, Aakash Bansal 等FSE 2024 · 被引用 15 次
- Modeling Hierarchical Syntax Structure with Triplet Position for Source Code SummarizationJuncai Guo, Jin Liu, Yao Wan, Li Li 等ACL 2022
- Attend, Translate and Summarize: An Efficient Method for Neural Cross-Lingual SummarizationJunnan Zhu, Yu Zhou, Jiajun Zhang, Chengqing ZongACL 2020 · 被引用 51 次
