Cross-language code search using static and dynamic analyses
George Mathew, Kathryn T. Stolee
Abstract
As code search permeates most activities in software development, code-to-code search has emerged to support using code as a query and retrieving similar code in the search results. Applications include duplicate code detection for refactoring, patch identification for program repair, and language translation. Existing code-to-code search tools rely on static similarity approaches such as the comparison of tokens and abstract syntax trees (AST) to approximate dynamic behavior, leading to low precision. Most tools do not support cross-language code-to-code search, and those that do, rely on machine learning models that require labeled training data.
We present Code-to-Code Search Across Languages (COSAL), a cross-language technique that uses both static and dynamic analyses to identify similar code and does not require a machine learning model. Code snippets are ranked using non-dominated sorting based on code token similarity, structural similarity, and behavioral similarity. We empirically evaluate COSAL on two datasets of 43,146 Java and Python files and 55,499 Java files and find that 1) code search based on non-dominated ranking of static and dynamic similarity measures is more effective compared to single or weighted measures; and 2) COSAL has better precision and recall compared to state-of-the-art within-language and cross-language code-tocode search tools. We explore the potential for using COSAL on large open-source repositories and discuss scalability to more languages and similarity metrics, providing a gateway for practical, multi-language code-to-code search.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ddddcb61-ba35-4500-8f04-cd3ce7ae9e02Cited by top-tier papers2
- CPVis: Evidence-based Multimodal Learning Analytics for Evaluation in Collaborative ProgrammingGefei Zhang, Shenming Ji, Yicao Li, Jingwei Tang et al.CHI 2025 · 8 citations
- Finding Compiler Bugs through Cross-Language Code Generator and Differential TestingQiong Feng, Xiaotian Ma, Ziyuan Feng, Marat Akhin et al.OOPSLA 2025 · 2 citations
Builds on3
- You Get Where You're Looking for: The Impact of Information Sources on Code SecurityYasemin Acar, Michael Backes, Sascha Fahl, Doowon Kim et al.S&P 2016 · 325 citations
- InferCode: Self-Supervised Learning of Code Representations by Predicting SubtreesNghi D. Q. Bui, Yijun Yu, Lingxiao JiangICSE 2021 · 106 citations
- SLACC: simion-based language agnostic code clonesGeorge Mathew, Chris Parnin, Kathryn T. StoleeICSE 2020 · 21 citations
Related papers
- Zero-Shot Cross-Domain Code Search without Fine-TuningKeyu Liang, Zhongxin Liu, Chao Liu, Zhiyuan Wan et al.FSE 2025 · 2 citations
- Accelerating Code Search with Deep Hashing and Code ClassificationWenchao Gu, Yanlin Wang, Lun Du, Hongyu Zhang et al.ACL 2022
- Multilingual Code Co-evolution using Large Language ModelsJiyang Zhang, Pengyu Nie, Junyi Jessy Li, Milos GligoricFSE 2023 · 34 citations
- CoCoSoDa: Effective Contrastive Learning for Code SearchEnsheng Shi, Yanlin Wang, Wenchao Gu, Lun Du et al.ICSE 2023 · 45 citations
- Semantic code search via equational reasoningVarot Premtoon, James Koppel, Armando Solar-LezamaPLDI 2020 · 49 citations
