ZC3: Zero-Shot Cross-Language Code Clone Detection
Jia Li, Chongyang Tao, Zhi Jin, Fang Liu, Jia Li, Ge Li
Abstract
Developers introduce code clones to improve programming productivity. Many existing studies have achieved impressive performance in monolingual code clone detection. However, during software development, more and more developers write semantically equivalent programs with different languages to support different platforms and help developers translate projects from one language to another. Considering that collecting cross-language parallel data, especially for low-resource languages, is expensive and time-consuming, how designing an effective cross-language model that does not rely on any parallel data is a significant problem. In this paper, we propose a novel method named ZC <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup> for Z_ero-shot Cross-language Code Clone detection. ZC <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup> designs the contrastive snippet prediction to form an isomorphic representation space among different programming languages. Based on this, ZC <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup> exploits domain-aware learning and cycle consistency learning to further constrain the model to generate representations that are aligned among different languages meanwhile are diacritical for different types of clones. To evaluate our approach, we conduct extensive experiments on four representative cross-language clone detection datasets. Experimental results show that ZC <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup> outperforms the state-of-the-art baselines by 67.12%, 51.39%, 14.85%, and 53.01% on the MAP score, respectively. We further investigate the representational distribution of different languages and discuss the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9308da5d-e6e0-4a4e-bcb9-cd9ba96350adCited by top-tier papers1
Ask how each one uses itBuilds on7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- EditSum: A Retrieve-and-Edit Framework for Source Code SummarizationJia Li, Yongmin Li, Ge Li, Xing Hu et al.ASE 2021 · 55 citations
- SkCoder: A Sketch-based Approach for Automatic Code GenerationJia Li, Yongmin Li, Ge Li, Zhi Jin et al.ICSE 2023 · 50 citations
- CODEP: Grammatical Seq2Seq Model for General-Purpose Code GenerationYihong Dong, Ge Li, Zhi JinISSTA 2023 · 18 citations
Related papers
- AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone DetectionYangkai Du, Tengfei Ma, Lingfei Wu, Xuhong Zhang et al.AAAI 2024 · 9 citations
- LC3: Long Cross-Language Code Clone Detection Enhanced by Opcode Sequences and Affinity AggregationXilin Lan, Huan Zhang, Yang Yang, Chengwu Xue et al.AAAI 2026
- Detecting Semantic Clones of Unseen FunctionalityKonstantinos Kitsios, Francesco Sovrano, Earl T. Barr, Alberto BacchelliASE 2025 · 1 citation
- The Struggles of LLMs in Cross-Lingual Code Clone DetectionMicheline Bénédicte Moumoula, Abdoul Kader Kaboré, Jacques Klein, Tegawendé F. BissyandéFSE 2025 · 4 citations
- CC2Vec: Combining Typed Tokens with Contrastive Learning for Effective Code Clone DetectionShihan Dou, Yueming Wu, Haoxiang Jia, Yuhao Zhou et al.FSE 2024 · 9 citations
