Gitor: Scalable Code Clone Detection by Building Global Sample Graph
Junjie Shan, Shihan Dou, Yueming Wu, Hairu Wu, Yang Liu
摘要
Code clone detection is about finding out similar code fragments, which has drawn much attention in software engineering since it is important for software maintenance and evolution. Researchers have proposed many techniques and tools for source code clone detection, but current detection methods concentrate on analyzing or processing code samples individually without exploring the underlying connections among code samples.
In this paper, we propose Gitor to capture the underlying connections among different code samples. Specifically, given a source code database, we first tokenize all code samples to extract the pre-defined individual information (e.g., keywords). After obtaining all samples' individual information, we leverage them to build a large global sample graph where each node is a code sample or a type of individual information. Then we apply a node embedding technique on the global sample graph to extract all the samples' vector representations. After collecting all code samples' vectors, we can simply compare the similarity between any two samples to detect possible clone pairs. More importantly, since the obtained vector of a sample is from a global sample graph, we can combine it with its own code features to improve the code clone detection performance. To demonstrate the effectiveness of Gitor, we evaluate it on a widely used dataset namely BigCloneBench. Our experimental results show that Gitor has higher accuracy in terms of code clone detection and excellent execution time for inputs of various sizes (1-100 MLOC) compared to existing state-of-the-art tools. Moreover, we also evaluate the combination of Gitor with other traditional vector-based clone detection methods, the results show that the use of Gitor enables them detect more code clones with higher F1.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Machine Learning is All You Need: A Simple Token-based Approach for Effective Code Clone DetectionSiyue Feng, Wenqi Suo, Yueming Wu, Deqing Zou 等ICSE 2024 · 被引用 20 次
- CC2Vec: Combining Typed Tokens with Contrastive Learning for Effective Code Clone DetectionShihan Dou, Yueming Wu, Haoxiang Jia, Yuhao Zhou 等FSE 2024 · 被引用 9 次
- One Bug, Hundreds Behind: LLMs for Large-Scale Bug DiscoveryQiushi Wu, Yue Xiao, Dhilung Kirat, Kevin Eykholt 等ICML 2026 · 被引用 5 次
- Feature Slice Matching for Precise Bug DetectionKe Ma, Jianjun Huang, Wei You, Bin Liang 等FSE 2026
它引用的顶会 Paper2
相关 Paper
- Tritor: Detecting Semantic Code Clones by Building Social Network-Based Triads ModelDeqing Zou, Siyue Feng, Yueming Wu, Wenqi Suo 等FSE 2023 · 被引用 6 次
- Detecting Semantic Code Clones by Building AST-based Markov Chains ModelYueming Wu, Siyue Feng, Deqing Zou, Hai JinASE 2022 · 被引用 21 次
- Fine-Grained Code Clone Detection with Block-Based Splitting of Abstract Syntax TreeTiancheng Hu, Zijing Xu, Yilin Fang, Yueming Wu 等ISSTA 2023 · 被引用 18 次
- CCGraph: a PDG-based code clone detector with approximate graph matchingYue Zou, Bihuan Ban, Yinxing Xue, Yun XuASE 2020 · 被引用 46 次
- SCDetector: Software Functional Clone Detection Based on Semantic Tokens AnalysisYueming Wu, Deqing Zou, Shihan Dou, Siru Yang 等ASE 2020 · 被引用 55 次
