Can Neural Clone Detection Generalize to Unseen Functionalitiesƒ
Chenyao Liu, Zeqi Lin, Jian-Guang Lou, Lijie Wen, Dongmei Zhang
摘要
Many recently proposed code clone detectors exploit neural networks to capture latent semantics of source code, thus achieving impressive results for detecting semantic clones. These neural clone detectors rely on the availability of large amounts of labeled training data. We identify a key oversight in the current evaluation methodology for neural clone detection: cross-functionality generalization (i.e., detecting semantic clones of which the functionalities are unseen in training). Specifically, we focus on this question: do neural clone detectors truly learn the ability to detect semantic clones, or they just learn how to model specific functionalities in training data while cannot generalize to realistic unseen functionalitiesƒ This paper investigates how the generalizability can be evaluated and improved.Our contributions are 3-folds: (1) We propose an evaluation methodology that can systematically measure the cross-functionality generalizability of neural clone detection. Based on this evaluation methodology, an empirical study is conducted and the results indicate that current neural clone detectors cannot generalize well as expected. (2) We conduct empirical analysis to understand key factors that can impact the generalizability. We investigate 3 factors: training data diversity, vocabulary, and locality. Results show that the performance loss on unseen functionalities can be reduced through addressing the out-of-vocabulary problem and increasing training data diversity. (3) We propose a human-in-the-loop mechanism that help adapt neural clone detectors to new code repositories containing lots of unseen functionalities. It improves annotation efficiency with the combination of transfer learning and active learning. Experimental results show that it reduces the amount of annotations by about 88%. Our code and data are publicly available <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> .
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- An empirical study of blockchain system vulnerabilities: modules, types, and patternsXiao Yi, Daoyuan Wu, Lingxiao Jiang, Yuzhou Fang 等FSE 2022 · 被引用 22 次
- Practical and Efficient Model Extraction of Sentiment Analysis APIsWeibin Wu, Jianping Zhang, Victor Junqiu Wei, Xixian Chen 等ICSE 2023 · 被引用 10 次
- Detecting Semantic Clones of Unseen FunctionalityKonstantinos Kitsios, Francesco Sovrano, Earl T. Barr, Alberto BacchelliASE 2025 · 被引用 1 次
相关 Paper
- Functional code clone detection with syntax and semantics fusion learningChunrong Fang, Zixi Liu, Yangyang Shi, Jeff Huang 等ISSTA 2020 · 被引用 125 次
- Tritor: Detecting Semantic Code Clones by Building Social Network-Based Triads ModelDeqing Zou, Siyue Feng, Yueming Wu, Wenqi Suo 等FSE 2023 · 被引用 6 次
- CC2Vec: Combining Typed Tokens with Contrastive Learning for Effective Code Clone DetectionShihan Dou, Yueming Wu, Haoxiang Jia, Yuhao Zhou 等FSE 2024 · 被引用 9 次
- AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone DetectionYangkai Du, Tengfei Ma, Lingfei Wu, Xuhong Zhang 等AAAI 2024 · 被引用 9 次
- The Struggles of LLMs in Cross-Lingual Code Clone DetectionMicheline Bénédicte Moumoula, Abdoul Kader Kaboré, Jacques Klein, Tegawendé F. BissyandéFSE 2025 · 被引用 4 次
