On the Importance of Building High-quality Training Datasets for Neural Code Search
Zhensu Sun, Li Li, Yan Liu, Xiaoning Du, Li Li
Abstract
The performance of neural code search is significantly influenced by the quality of the training data from which the neural models are derived. A large corpus of high-quality query and code pairs is demanded to establish a precise mapping from the natural language to the programming language. Due to the limited availability, most widely-used code search datasets are established with compromise, such as using code comments as a replacement of queries. Our empirical study on a famous code search dataset reveals that over one-third of its queries contain noises that make them deviate from natural user queries. Models trained through noisy data are faced with severe performance degradation when applied in real-world scenarios. To improve the dataset quality and make the queries of its samples semantically identical to real user queries is critical for the practical usability of neural code search. In this paper, we propose a data cleaning framework consisting of two subsequent filters: a rule-based syntactic filter and a model-based semantic filter. This is the first framework that applies semantic query cleaning to code search datasets. Experimentally, we evaluated the effectiveness of our framework on two widely-used code search models and three manually-annotated code retrieval benchmarks. Training the popular DeepCS model with the filtered dataset from our framework improves its performance by 19.2% MRR and 21.3% Answer@1, on average with the three validation benchmarks. CCS CONCEPTS • Software and its engineering → Reusability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb245b43-586c-4cdc-9484-061f36fa4055Cited by top-tier papers20
- No more fine-tuning? an experimental evaluation of prompt tuning in code intelligenceChaozheng Wang, Yuanhang Yang, Cuiyun Gao, Yun Peng et al.FSE 2022 · 148 citations
- Data Quality for Software Vulnerability DatasetsRoland Croft, Muhammad Ali Babar, M. Mehdi KholoosiICSE 2023 · 138 citations
- CCT5: A Code-Change-Oriented Pre-trained ModelBo Lin, Shangwen Wang, Zhongxin Liu, Yepang Liu et al.FSE 2023 · 69 citations
- Are we building on the rock? on the importance of data preprocessing for code summarizationLin Shi, Fangwen Mu, Xiao Chen, Song Wang et al.FSE 2022 · 65 citations
- LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and MitigationZiyao Zhang, Chong Wang, Yanlin Wang, Ensheng Shi et al.ISSTA 2025 · 53 citations
Builds on1
Related papers
- You see what I want you to see: poisoning vulnerabilities in neural code searchYao Wan, Shijie Zhang, Hongyu Zhang, Yulei Sui et al.FSE 2022 · 57 citations
- On-the-fly Improving Performance of Deep Code Models via Input DenoisingZhao Tian, Junjie Chen, Xiangyu ZhangASE 2023 · 8 citations
- Less Is More: On the Importance of Data Quality for Unit Test GenerationJunwei Zhang, Xing Hu, Shan Gao, Xin Xia et al.FSE 2025 · 2 citations
- Exploring Representation-level Augmentation for Code SearchHaochen Li, Chunyan Miao, Cyril Leung, Yanxian Huang et al.EMNLP 2022 · 17 citations
- LLM-Assisted Code Cleaning For Training Accurate Code GeneratorsNaman Jain, Tianjun Zhang, Wei-Lin Chiang, Joseph E. Gonzalez et al.ICLR 2024 · 49 citations
