Are we building on the rock? on the importance of data preprocessing for code summarization
Lin Shi, Fangwen Mu, Xiao Chen, Song Wang, Junjie Wang, Ye Yang, Ge Li, Xin Xia, Qing Wang
Abstract
Code summarization, the task of generating useful comments given the code, has long been of interest. Most of the existing code summarization models are trained and validated on widely-used code comment benchmark datasets. However, little is known about the quality of the benchmark datasets built from real-world projects. Are the benchmark datasets as good as expected? To bridge the gap, we conduct a systematic research to assess and improve the quality of four benchmark datasets widely used for code summarization tasks. First, we propose an automated code-comment cleaning tool that can accurately detect noisy data caused by inappropriate data preprocessing operations from existing benchmark datasets. Then, we apply the tool to further assess the data quality of the four benchmark datasets, based on the detected noises. Finally, we conduct comparative experiments to investigate the impact of noisy data on the performance of code summarization models. The results show that these data preprocessing noises widely exist in all four benchmark datasets, and removing these noisy data leads to a significant improvement on the performance of code summarization. We believe that the findings and insights will enable a better understanding of data quality in code summarization tasks, and pave the way for relevant research and practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58f4572d-e01b-4ea9-9c00-45e4f94cf653Cited by top-tier papers10
- Data Quality for Software Vulnerability DatasetsRoland Croft, Muhammad Ali Babar, M. Mehdi KholoosiICSE 2023 · 138 citations
- Developer-Intent Driven Code Comment GenerationFangwen Mu, Xiao Chen, Lin Shi, Song Wang et al.ICSE 2023 · 25 citations
- Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code ModelsShuzheng Gao, Wenxin Mao, Cuiyun Gao, Li Li et al.ICSE 2024 · 15 citations
- Keeping Pace with Ever-Increasing Data: Towards Continual Learning of Code Intelligence ModelsShuzheng Gao, Hongyu Zhang, Cuiyun Gao, Chaozheng WangICSE 2023 · 14 citations
- SimLLM: Calculating Semantic Similarity in Code Summaries using a Large Language Model-Based ApproachXin Jin, Zhiqiang LinFSE 2024 · 8 citations
Builds on12
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Retrieval-based neural source code summarizationJian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun et al.ICSE 2020 · 242 citations
- Reassessing automatic evaluation metrics for code summarization tasksDevjeet Roy, Sarah Fakhoury, Venera ArnaoudovaFSE 2021 · 103 citations
- Leveraging Code Generation to Improve Code Retrieval and Summarization via Dual LearningWei Ye, Rui Xie, Jinglei Zhang, Tianxiang Hu et al.WWW 2020 · 83 citations
- On the Evaluation of Neural Code SummarizationEnsheng Shi, Yanlin Wang, Lun Du, Junjie Chen et al.ICSE 2022 · 76 citations
Related papers
- On the Importance of Building High-quality Training Datasets for Neural Code SearchZhensu Sun, Li Li, Yan Liu, Xiaoning Du et al.ICSE 2022 · 67 citations
- Intrinsic Evaluation of Summarization DatasetsRishi Bommasani, Claire CardieEMNLP 2020 · 52 citations
- Impact of Evaluation Methodologies on Code SummarizationPengyu Nie, Jiyang Zhang, Junyi Jessy Li, Raymond J. Mooney et al.ACL 2022 · 21 citations
- Automating code review activities by large-scale pre-trainingZhiyu Li, Shuai Lu, Daya Guo, Nan Duan et al.FSE 2022 · 195 citations
- How Effectively Do Code Language Models Understand Poor-Readability Code?Chao Hu, Yitian Chai, Hao Zhou, Fandong Meng et al.ASE 2024 · 3 citations
