Self-Supervised Contrastive Learning for Code Retrieval and Summarization via Semantic-Preserving Transformations
Nghi D. Q. Bui, Yijun Yu, Lingxiao Jiang
Abstract
We propose Corder, a self-supervised contrastive learning framework for source code model. Corder is designed to alleviate the need of labeled data for code retrieval and code summarization tasks. The pre-trained model of Corder can be used in two ways: (1) it can produce vector representation of code which can be applied to code retrieval tasks that do not have labeled data; (2) it can be used in a fine-tuning process for tasks that might still require label data such as code summarization. The key innovation is that we train the source code model by asking it to recognize similar and dissimilar code snippets through a contrastive learning objective. To do so, we use a set of semantic-preserving transformation operators to generate code snippets that are syntactically diverse but semantically equivalent. Through extensive experiments, we have shown that the code models pretrained by Corder substantially outperform the other baselines for code-to-code retrieval, text-to-code retrieval, and code-to-text summarization tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6496f543-0112-4ec8-9632-3ad5272e85d7Cited by top-tier papers38
- ReACC: A Retrieval-Augmented Code Completion FrameworkShuai Lu, Nan Duan, Hojae Han, Daya Guo et al.ACL 2022 · 208 citations
- Path-sensitive code embedding via contrastive learning for software vulnerability detectionXiao Cheng, Guanqin Zhang, Haoyu Wang, Yulei SuiISSTA 2022 · 98 citations
- SelfAPR: Self-supervised Program Repair with Test Execution DiagnosticsHe Ye, Matias Martinez, Xiapu Luo, Tao Zhang et al.ASE 2022 · 75 citations
- Bridging Pre-trained Models and Downstream Tasks for Source Code UnderstandingDeze Wang, Zhouyang Jia, Shanshan Li, Yue Yu et al.ICSE 2022 · 68 citations
- ContraBERT: Enhancing Code Pre-trained Models via Contrastive LearningShangqing Liu, Bozhi Wu, Xiaofei Xie, Guozhu Meng et al.ICSE 2023 · 56 citations
Builds on11
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Graph-based, Self-Supervised Program Repair from Diagnostic FeedbackMichihiro Yasunaga, Percy LiangICML 2020 · 198 citations
- A Mutual Information Maximization Perspective of Language Representation LearningLingpeng Kong, Cyprien de Masson d'Autume, Lei Yu, Wang Ling et al.ICLR 2020 · 179 citations
- Generating Adversarial Examples for Holding Robustness of Source Code Processing ModelsHuangzhao Zhang, Zhuo Li, Ge Li, Lei Ma et al.AAAI 2020 · 148 citations
- Structural Language Models of CodeUri Alon, Roy Sadaka, Omer Levy, Eran YahavICML 2020 · 115 citations
Related papers
- CodeRetriever: A Large Scale Contrastive Pre-Training Method for Code SearchXiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu et al.EMNLP 2022 · 25 citations
- Towards Better Code Understanding in Decoder-Only Models with Contrastive LearningJiayi Lin, Yanlin Wang, Yibiao Yang, Lei Zhang et al.AAAI 2026 · 2 citations
- UniCoR: Modality Collaboration for Robust Cross-Language Hybrid Code RetrievalYang Yang, Li Kuang, Jiakun Liu, Zhongxin Liu et al.ICSE 2026
- Contrastive Code Representation LearningParas Jain, Ajay Jain, Tianjun Zhang, Pieter Abbeel et al.EMNLP 2021 · 11 citations
- CoCoSoDa: Effective Contrastive Learning for Code SearchEnsheng Shi, Yanlin Wang, Wenchao Gu, Lun Du et al.ICSE 2023 · 45 citations
