CoCoSoDa: Effective Contrastive Learning for Code Search
Ensheng Shi, Yanlin Wang, Wenchao Gu, Lun Du, Hongyu Zhang, Shi Han, Dongmei Zhang, Hongbin Sun
Abstract
Code search aims to retrieve semantically relevant code snippets for a given natural language query. Recently, many approaches employing contrastive learning have shown promising results on code representation learning and greatly improved the performance of code search. However, there is still a lot of room for improvement in using contrastive learning for code search. In this paper, we propose CoCoSoDa to effectively utilize contrastive learning for code search via two key factors in contrastive learning: data augmentation and negative samples. Specifically, soft data augmentation is to dynamically masking or replacing some tokens with their types for input sequences to generate positive samples. Momentum mechanism is used to generate large and consistent representations of negative samples in a mini-batch through maintaining a queue and a momentum encoder. In addition, multimodal contrastive learning is used to pull together representations of code-query pairs and push apart the unpaired code snippets and queries. We conduct extensive experiments to evaluate the effectiveness of our approach on a large-scale dataset with six programming languages. Experimental results show that: (1) CoCoSoDa outperforms 18 baselines and especially exceeds CodeBERT, GraphCodeBERT, and UniXcoder by 13.3%, 10.5%, and 5.9% on average MRR scores, respectively. (2) The ablation studies show the effectiveness of each component of our approach. (3) We adapt our techniques to several different pre-trained models such as RoBERTa, CodeBERT, and GraphCodeBERT and observe a significant boost in their performance in code search. (4) Our model performs robustly under different hyper-parameters. Furthermore, we perform qualitative and quantitative analyses to explore reasons behind the good performance of our model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- What Makes Good In-Context Demonstrations for Code Intelligence Tasks with LLMs?Shuzheng Gao, Xin-Cheng Wen, Cuiyun Gao, Wenxuan Wang et al.ASE 2023 · 80 citations
- Two Birds with One Stone: Boosting Code Generation and Code Search via a Generative Adversarial NetworkShangwen Wang, Bo Lin, Zhensu Sun, Ming Wen et al.OOPSLA 2023 · 21 citations
- Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code ModelsShuzheng Gao, Wenxin Mao, Cuiyun Gao, Li Li et al.ICSE 2024 · 15 citations
- Instructive Code Retriever: Learn from Large Language Model's Feedback for Code Intelligence TasksJiawei Lu, Haoye Wang, Zhongxin Liu, Keyu Liang et al.ASE 2024 · 3 citations
- SECRET: Towards Scalable and Efficient Code Retrieval via Segmented Deep HashingWenchao Gu, Ensheng Shi, Yanlin Wang, Lun Du et al.ICSE 2025 · 3 citations
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
Related papers
- Exploring Representation-level Augmentation for Code SearchHaochen Li, Chunyan Miao, Cyril Leung, Yanxian Huang et al.EMNLP 2022 · 17 citations
- Uncertainty-Aware Contrastive Learning with Hard Negative Sampling for Code Search TasksHan Liu, Jiaqing Zhan, Qin ZhangAAAI 2025 · 1 citation
- ContraBERT: Enhancing Code Pre-trained Models via Contrastive LearningShangqing Liu, Bozhi Wu, Xiaofei Xie, Guozhu Meng et al.ICSE 2023 · 56 citations
- CodeRetriever: A Large Scale Contrastive Pre-Training Method for Code SearchXiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu et al.EMNLP 2022 · 25 citations
- Self-Supervised Contrastive Learning for Code Retrieval and Summarization via Semantic-Preserving TransformationsNghi D. Q. Bui, Yijun Yu, Lingxiao JiangSIGIR 2021 · 98 citations
