On the Importance of Building High-quality Training Datasets for Neural Code Search
Zhensu Sun, Li Li, Yan Liu, Xiaoning Du, Li Li
摘要
The performance of neural code search is significantly influenced by the quality of the training data from which the neural models are derived. A large corpus of high-quality query and code pairs is demanded to establish a precise mapping from the natural language to the programming language. Due to the limited availability, most widely-used code search datasets are established with compromise, such as using code comments as a replacement of queries. Our empirical study on a famous code search dataset reveals that over one-third of its queries contain noises that make them deviate from natural user queries. Models trained through noisy data are faced with severe performance degradation when applied in real-world scenarios. To improve the dataset quality and make the queries of its samples semantically identical to real user queries is critical for the practical usability of neural code search. In this paper, we propose a data cleaning framework consisting of two subsequent filters: a rule-based syntactic filter and a model-based semantic filter. This is the first framework that applies semantic query cleaning to code search datasets. Experimentally, we evaluated the effectiveness of our framework on two widely-used code search models and three manually-annotated code retrieval benchmarks. Training the popular DeepCS model with the filtered dataset from our framework improves its performance by 19.2% MRR and 21.3% Answer@1, on average with the three validation benchmarks. CCS CONCEPTS • Software and its engineering → Reusability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- No more fine-tuning? an experimental evaluation of prompt tuning in code intelligenceChaozheng Wang, Yuanhang Yang, Cuiyun Gao, Yun Peng 等FSE 2022 · 被引用 148 次
- Data Quality for Software Vulnerability DatasetsRoland Croft, Muhammad Ali Babar, M. Mehdi KholoosiICSE 2023 · 被引用 138 次
- CCT5: A Code-Change-Oriented Pre-trained ModelBo Lin, Shangwen Wang, Zhongxin Liu, Yepang Liu 等FSE 2023 · 被引用 69 次
- Are we building on the rock? on the importance of data preprocessing for code summarizationLin Shi, Fangwen Mu, Xiao Chen, Song Wang 等FSE 2022 · 被引用 65 次
- LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and MitigationZiyao Zhang, Chong Wang, Yanlin Wang, Ensheng Shi 等ISSTA 2025 · 被引用 53 次
它引用的顶会 Paper1
相关 Paper
- You see what I want you to see: poisoning vulnerabilities in neural code searchYao Wan, Shijie Zhang, Hongyu Zhang, Yulei Sui 等FSE 2022 · 被引用 57 次
- On-the-fly Improving Performance of Deep Code Models via Input DenoisingZhao Tian, Junjie Chen, Xiangyu ZhangASE 2023 · 被引用 8 次
- Less Is More: On the Importance of Data Quality for Unit Test GenerationJunwei Zhang, Xing Hu, Shan Gao, Xin Xia 等FSE 2025 · 被引用 2 次
- Exploring Representation-level Augmentation for Code SearchHaochen Li, Chunyan Miao, Cyril Leung, Yanxian Huang 等EMNLP 2022 · 被引用 17 次
- LLM-Assisted Code Cleaning For Training Accurate Code GeneratorsNaman Jain, Tianjun Zhang, Wei-Lin Chiang, Joseph E. Gonzalez 等ICLR 2024 · 被引用 49 次
