Can Neural Clone Detection Generalize to Unseen Functionalitiesƒ
Chenyao Liu, Zeqi Lin, Jian-Guang Lou, Lijie Wen, Dongmei Zhang
Abstract
Many recently proposed code clone detectors exploit neural networks to capture latent semantics of source code, thus achieving impressive results for detecting semantic clones. These neural clone detectors rely on the availability of large amounts of labeled training data. We identify a key oversight in the current evaluation methodology for neural clone detection: cross-functionality generalization (i.e., detecting semantic clones of which the functionalities are unseen in training). Specifically, we focus on this question: do neural clone detectors truly learn the ability to detect semantic clones, or they just learn how to model specific functionalities in training data while cannot generalize to realistic unseen functionalitiesƒ This paper investigates how the generalizability can be evaluated and improved.Our contributions are 3-folds: (1) We propose an evaluation methodology that can systematically measure the cross-functionality generalizability of neural clone detection. Based on this evaluation methodology, an empirical study is conducted and the results indicate that current neural clone detectors cannot generalize well as expected. (2) We conduct empirical analysis to understand key factors that can impact the generalizability. We investigate 3 factors: training data diversity, vocabulary, and locality. Results show that the performance loss on unseen functionalities can be reduced through addressing the out-of-vocabulary problem and increasing training data diversity. (3) We propose a human-in-the-loop mechanism that help adapt neural clone detectors to new code repositories containing lots of unseen functionalities. It improves annotation efficiency with the combination of transfer learning and active learning. Experimental results show that it reduces the amount of annotations by about 88%. Our code and data are publicly available <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> .
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- An empirical study of blockchain system vulnerabilities: modules, types, and patternsXiao Yi, Daoyuan Wu, Lingxiao Jiang, Yuzhou Fang et al.FSE 2022 · 22 citations
- Practical and Efficient Model Extraction of Sentiment Analysis APIsWeibin Wu, Jianping Zhang, Victor Junqiu Wei, Xixian Chen et al.ICSE 2023 · 10 citations
- Detecting Semantic Clones of Unseen FunctionalityKonstantinos Kitsios, Francesco Sovrano, Earl T. Barr, Alberto BacchelliASE 2025 · 1 citation
Related papers
- Functional code clone detection with syntax and semantics fusion learningChunrong Fang, Zixi Liu, Yangyang Shi, Jeff Huang et al.ISSTA 2020 · 125 citations
- Tritor: Detecting Semantic Code Clones by Building Social Network-Based Triads ModelDeqing Zou, Siyue Feng, Yueming Wu, Wenqi Suo et al.FSE 2023 · 6 citations
- CC2Vec: Combining Typed Tokens with Contrastive Learning for Effective Code Clone DetectionShihan Dou, Yueming Wu, Haoxiang Jia, Yuhao Zhou et al.FSE 2024 · 9 citations
- AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone DetectionYangkai Du, Tengfei Ma, Lingfei Wu, Xuhong Zhang et al.AAAI 2024 · 9 citations
- The Struggles of LLMs in Cross-Lingual Code Clone DetectionMicheline Bénédicte Moumoula, Abdoul Kader Kaboré, Jacques Klein, Tegawendé F. BissyandéFSE 2025 · 4 citations
