NIL: large-scale detection of large-variance clones
Tasuku Nakagawa, Yoshiki Higo, Shinji Kusumoto
Abstract
A code clone (in short, clone) is a code fragment that is identical or similar to other code fragments in source code. Clones generated by a large number of changes to copy-and-pasted code fragments are called large-variance (modifications are scattered) or large-gap (modifications are in one place) clones. It is difficult for general clone detection techniques to detect such clones and thus specialized techniques are necessary. In addition, with the rapid growth of software development, scalable clone detectors that can detect clones in large codebases are required. However, there are no existing techniques for quickly detecting large-variance or large-gap clones in large codebases. In this paper, we propose a scalable clone detection technique that can detect large-variance clones from large codebases and describe its implementation, called NIL. NIL is a token-based clone detector that efficiently identifies clone candidates using an N-gram representation of token sequences and an inverted index. Then, NIL verifies the clone candidates by measuring their similarity based on the longest common subsequence between their token sequences. We evaluate NIL in terms of large- variance clone detection accuracy, general Type-1, Type-2, and Type- 3 clone detection accuracy, and scalability. Our experimental results show that NIL has higher accuracy in terms of large-variance clone detection, equivalent accuracy in terms of general clone detection, and the shortest execution time for inputs of various sizes (1–250 MLOC) compared to existing state-of-the-art tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 865b9a86-aaaf-459c-aaa3-5b64a4c7dbc4Cited by top-tier papers12
- OSSFP: Precise and Scalable C/C++ Third-Party Library Detection using Fingerprinting FunctionsJiahui Wu, Zhengzi Xu, Wei Tang, Lyuye Zhang et al.ICSE 2023 · 29 citations
- Learning Graph-based Code Representations for Source-level Functional Similarity DetectionJiahao Liu, Jun Zeng, Xiang Wang, Zhenkai LiangICSE 2023 · 27 citations
- Machine Learning is All You Need: A Simple Token-based Approach for Effective Code Clone DetectionSiyue Feng, Wenqi Suo, Yueming Wu, Deqing Zou et al.ICSE 2024 · 20 citations
- Fine-Grained Code Clone Detection with Block-Based Splitting of Abstract Syntax TreeTiancheng Hu, Zijing Xu, Yilin Fang, Yueming Wu et al.ISSTA 2023 · 18 citations
- Comparison and Evaluation of Clone Detection Techniques with Different Code RepresentationsYuekun Wang, Yuhang Ye, Yueming Wu, Weiwei Zhang et al.ICSE 2023 · 15 citations
Builds on1
Related papers
- CC2Vec: Combining Typed Tokens with Contrastive Learning for Effective Code Clone DetectionShihan Dou, Yueming Wu, Haoxiang Jia, Yuhao Zhou et al.FSE 2024 · 9 citations
- Detecting Semantic Code Clones by Building AST-based Markov Chains ModelYueming Wu, Siyue Feng, Deqing Zou, Hai JinASE 2022 · 21 citations
- Tritor: Detecting Semantic Code Clones by Building Social Network-Based Triads ModelDeqing Zou, Siyue Feng, Yueming Wu, Wenqi Suo et al.FSE 2023 · 6 citations
- Gitor: Scalable Code Clone Detection by Building Global Sample GraphJunjie Shan, Shihan Dou, Yueming Wu, Hairu Wu et al.FSE 2023 · 7 citations
- TreeCen: Building Tree Graph for Scalable Semantic Code Clone DetectionYutao Hu, Deqing Zou, Junru Peng, Yueming Wu et al.ASE 2022 · 30 citations
