Retrieval-Augmented Generation for Code Summarization via Hybrid GNN
Shangqing Liu, Yu Chen, Xiaofei Xie, Jing Kai Siow, Yang Liu
Abstract
Source code summarization aims to generate natural language summaries from structured code snippets for better understanding code functionalities. However, automatic code summarization is challenging due to the complexity of the source code and the language gap between the source code and natural language summaries. Most previous approaches either rely on retrieval-based (which can take advantage of similar examples seen from the retrieval database, but have low generalization performance) or generation-based methods (which have better generalization performance, but cannot take advantage of similar examples). This paper proposes a novel retrieval-augmented mechanism to combine the benefits of both worlds. Furthermore, to mitigate the limitation of Graph Neural Networks (GNNs) on capturing global graph structure information of source code, we propose a novel attention-based dynamic graph to complement the static graph representation of the source code, and design a hybrid message passing GNN for capturing both the local and global structural information. To evaluate the proposed approach, we release a new challenging benchmark, crawled from diversified large-scale open-source C projects (total 95k+ unique functions in the dataset). Our method achieves the state-of-the-art performance, improving existing methods by 1.42, 2.44 and 1.29 in terms of BLEU-4, ROUGE-L and METEOR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de7d9685-d2dd-44b8-b721-03e6c8ce5319Cited by top-tier papers36
- No more fine-tuning? an experimental evaluation of prompt tuning in code intelligenceChaozheng Wang, Yuanhang Yang, Cuiyun Gao, Yun Peng et al.FSE 2022 · 148 citations
- NatGen: generative pre-training by "naturalizing" source codeSaikat Chakraborty, Toufique Ahmed, Yangruibo Ding, Premkumar T. Devanbu et al.FSE 2022 · 101 citations
- ASIF: Coupled Data Turns Unimodal Models to Multimodal without TrainingAntonio Norelli, Marco Fumero, Valentino Maiorca, Luca Moschella et al.NeurIPS 2023 · 57 citations
- ContraBERT: Enhancing Code Pre-trained Models via Contrastive LearningShangqing Liu, Bozhi Wu, Xiaofei Xie, Guozhu Meng et al.ICSE 2023 · 56 citations
- Automated Assertion Generation via Information Retrieval and Its Integration with Deep learningHao Yu, Yiling Lou, Ke Sun, Dezhi Ran et al.ICSE 2022 · 42 citations
Builds on6
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 1,586 citations
- Measuring and Relieving the Over-Smoothing Problem for Graph Neural Networks from the Topological ViewDeli Chen, Yankai Lin, Wei Li, Peng Li et al.AAAI 2020 · 1,353 citations
- PairNorm: Tackling Oversmoothing in GNNsLingxiao Zhao, Leman AkogluICLR 2020 · 590 citations
- Iterative Deep Graph Learning for Graph Neural Networks: Better and Robust Node EmbeddingsYu Chen, Lingfei Wu, Mohammed J. ZakiNeurIPS 2020 · 559 citations
- Retrieval-based neural source code summarizationJian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun et al.ICSE 2020 · 242 citations
Related papers
- Modeling Hierarchical Syntax Structure with Triplet Position for Source Code SummarizationJuncai Guo, Jin Liu, Yao Wan, Li Li et al.ACL 2022
- MGF-ESE: An Enhanced Semantic Extractor with Multi-Granularity Feature Fusion for Code SummarizationXiaolong Xu, Yuxin Cao, Hongsheng Hu, Haolong Xiang et al.WWW 2025 · 4 citations
- CAST: Enhancing Code Summarization with Hierarchical Splitting and Reconstruction of Abstract Syntax TreesEnsheng Shi, Yanlin Wang, Lun Du, Hongyu Zhang et al.EMNLP 2021 · 42 citations
- EditSum: A Retrieve-and-Edit Framework for Source Code SummarizationJia Li, Yongmin Li, Ge Li, Xing Hu et al.ASE 2021 · 55 citations
- Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and BeyondMinh Le-Anh, Huyen Nguyen, Khanh An Tran, Nam Le Hai et al.FSE 2026
