CONCORD: Clone-Aware Contrastive Learning for Source Code
Yangruibo Ding, Saikat Chakraborty, Luca Buratti, Saurabh Pujar, Alessandro Morari, Gail E. Kaiser, Baishakhi Ray
Abstract
Deep Learning (DL) models to analyze source code have shown immense promise during the past few years. More recently, selfsupervised pre-training has gained traction for learning generic code representations valuable for many downstream SE tasks, such as clone and bug detection. While previous work successfully learned from different code abstractions (e.g., token, AST, graph), we argue that it is also essential to factor in how developers code day-to-day for learning generalpurpose representation. On the one hand, human developers tend to write repetitive programs referencing existing code snippets from the current codebase or online resources (e.g., Stack Overflow website) rather than implementing functions from scratch; such behaviors result in a vast number of code clones. In contrast, a deviant clone by mistake might trigger malicious program behaviors. Thus, as a proxy to incorporate developers' coding behavior into the pre-training scheme, we propose to include code clones and their deviants. In particular, we propose CONCORD, a self-supervised pre-training strategy to place benign clones closer in the representation space while moving deviants further apart. We show that CONCORD's clone-aware pre-training drastically reduces the need for expensive pre-training resources while improving the performance of downstream SE tasks. We also empirically demonstrate that CONCORD can improve existing pre-trained models to learn better representations that consequently become more efficient in both identifying semantically equivalent programs and differentiating buggy from non-buggy code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1318b76b-537e-46f8-872b-1cfbbe6a0fe0Cited by top-tier papers4
- Vulnerability Detection with Code Language Models: How Far are We?Yangruibo Ding, Yanjun Fu, Omniyyah Ibrahim, Chawin Sitawarin et al.ICSE 2025 · 44 citations
- Coca: Improving and Explaining Graph Neural Network-Based Vulnerability Detection SystemsSicong Cao, Xiaobing Sun, Xiaoxue Wu, David Lo et al.ICSE 2024 · 27 citations
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller et al.NDSS 2026 · 17 citations
- Feature Slice Matching for Precise Bug DetectionKe Ma, Jianjun Huang, Wei You, Bin Liang et al.FSE 2026
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
Related papers
- Towards Learning (Dis)-Similarity of Source Code from Program ContrastsYangruibo Ding, Luca Buratti, Saurabh Pujar, Alessandro Morari et al.ACL 2022
- Contrastive Code Representation LearningParas Jain, Ajay Jain, Tianjun Zhang, Pieter Abbeel et al.EMNLP 2021 · 11 citations
- InferCode: Self-Supervised Learning of Code Representations by Predicting SubtreesNghi D. Q. Bui, Yijun Yu, Lingxiao JiangICSE 2021 · 106 citations
- AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone DetectionYangkai Du, Tengfei Ma, Lingfei Wu, Xuhong Zhang et al.AAAI 2024 · 9 citations
- Bridging Pre-trained Models and Downstream Tasks for Source Code UnderstandingDeze Wang, Zhouyang Jia, Shanshan Li, Yue Yu et al.ICSE 2022 · 68 citations
