Are we building on the rock? on the importance of data preprocessing for code summarization
Lin Shi, Fangwen Mu, Xiao Chen, Song Wang, Junjie Wang, Ye Yang, Ge Li, Xin Xia, Qing Wang
摘要
Code summarization, the task of generating useful comments given the code, has long been of interest. Most of the existing code summarization models are trained and validated on widely-used code comment benchmark datasets. However, little is known about the quality of the benchmark datasets built from real-world projects. Are the benchmark datasets as good as expected? To bridge the gap, we conduct a systematic research to assess and improve the quality of four benchmark datasets widely used for code summarization tasks. First, we propose an automated code-comment cleaning tool that can accurately detect noisy data caused by inappropriate data preprocessing operations from existing benchmark datasets. Then, we apply the tool to further assess the data quality of the four benchmark datasets, based on the detected noises. Finally, we conduct comparative experiments to investigate the impact of noisy data on the performance of code summarization models. The results show that these data preprocessing noises widely exist in all four benchmark datasets, and removing these noisy data leads to a significant improvement on the performance of code summarization. We believe that the findings and insights will enable a better understanding of data quality in code summarization tasks, and pave the way for relevant research and practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Data Quality for Software Vulnerability DatasetsRoland Croft, Muhammad Ali Babar, M. Mehdi KholoosiICSE 2023 · 被引用 138 次
- Developer-Intent Driven Code Comment GenerationFangwen Mu, Xiao Chen, Lin Shi, Song Wang 等ICSE 2023 · 被引用 25 次
- Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code ModelsShuzheng Gao, Wenxin Mao, Cuiyun Gao, Li Li 等ICSE 2024 · 被引用 15 次
- Keeping Pace with Ever-Increasing Data: Towards Continual Learning of Code Intelligence ModelsShuzheng Gao, Hongyu Zhang, Cuiyun Gao, Chaozheng WangICSE 2023 · 被引用 14 次
- SimLLM: Calculating Semantic Similarity in Code Summaries using a Large Language Model-Based ApproachXin Jin, Zhiqiang LinFSE 2024 · 被引用 8 次
它引用的顶会 Paper12
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Retrieval-based neural source code summarizationJian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun 等ICSE 2020 · 被引用 242 次
- Reassessing automatic evaluation metrics for code summarization tasksDevjeet Roy, Sarah Fakhoury, Venera ArnaoudovaFSE 2021 · 被引用 103 次
- Leveraging Code Generation to Improve Code Retrieval and Summarization via Dual LearningWei Ye, Rui Xie, Jinglei Zhang, Tianxiang Hu 等WWW 2020 · 被引用 83 次
- On the Evaluation of Neural Code SummarizationEnsheng Shi, Yanlin Wang, Lun Du, Junjie Chen 等ICSE 2022 · 被引用 76 次
相关 Paper
- On the Importance of Building High-quality Training Datasets for Neural Code SearchZhensu Sun, Li Li, Yan Liu, Xiaoning Du 等ICSE 2022 · 被引用 67 次
- Intrinsic Evaluation of Summarization DatasetsRishi Bommasani, Claire CardieEMNLP 2020 · 被引用 52 次
- Impact of Evaluation Methodologies on Code SummarizationPengyu Nie, Jiyang Zhang, Junyi Jessy Li, Raymond J. Mooney 等ACL 2022 · 被引用 21 次
- Automating code review activities by large-scale pre-trainingZhiyu Li, Shuai Lu, Daya Guo, Nan Duan 等FSE 2022 · 被引用 195 次
- How Effectively Do Code Language Models Understand Poor-Readability Code?Chao Hu, Yitian Chai, Hao Zhou, Fandong Meng 等ASE 2024 · 被引用 3 次
