Uncertainty-Aware Contrastive Learning with Hard Negative Sampling for Code Search Tasks
Han Liu, Jiaqing Zhan, Qin Zhang
摘要
Code search is a highly required technique for software development. In recent years, the rapid development of transformer-based language models has made it increasingly more popular to adapt a pre-trained language model to a code search task, where contrastive learning is typically adopted to semantically align user queries and codes in an embedding space. Considering that the same semantic meaning can be presented using diverse language styles in user queries and codes, the representation of queries and codes in an embedding space may thus be non-deterministic. To address the above-specified point, this paper proposes an uncertainty-aware contrastive learning approach for code search. Specifically, for both queries and codes, we design an uncertainty learning strategy to produce diverse embeddings by learning to transform the original inputs into Gaussian distributions and then taking a reparameterization trick. We also design a hard negative sampling strategy to construct query-code pairs for improving the effectiveness of uncertainty-aware contrastive learning. The experimental results indicate that our approach outperforms 10 baseline methods on a large code search dataset with six programming languages. The results also show that our strategies of uncertainty learning and hard negative sampling can really help enhance the representation of queries and codes leading to an improvement of the code search performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SNAPHARD CONTRAST LEARNINGChangpu Meng, Jie Yang, Wanqing Li, Yi GuoICLR 2026
- UniCoR: Modality Collaboration for Robust Cross-Language Hybrid Code RetrievalYang Yang, Li Kuang, Jiakun Liu, Zhongxin Liu 等ICSE 2026
它引用的顶会 Paper16
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Robust Person Re-Identification by Modelling Feature UncertaintyTianyuan Yu, Da Li, Yongxin Yang, Timothy M. Hospedales 等ICCV 2019 · 被引用 148 次
相关 Paper
- Exploring Representation-level Augmentation for Code SearchHaochen Li, Chunyan Miao, Cyril Leung, Yanxian Huang 等EMNLP 2022 · 被引用 17 次
- CoCoSoDa: Effective Contrastive Learning for Code SearchEnsheng Shi, Yanlin Wang, Wenchao Gu, Lun Du 等ICSE 2023 · 被引用 45 次
- Code Representation Learning at ScaleDejiao Zhang, Wasi Uddin Ahmad, Ming Tan, Hantian Ding 等ICLR 2024 · 被引用 32 次
- CodeRetriever: A Large Scale Contrastive Pre-Training Method for Code SearchXiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu 等EMNLP 2022 · 被引用 25 次
- Hedgecode: A Multi-Task Hedging Contrastive Learning Framework for Code SearchGong Chen, Xiaoyuan Xie, Daniel Tang, Qi Xin 等ICSE 2025
