KEENHash: Hashing Programs into Function-Aware Embeddings for Large-Scale Binary Code Similarity Analysis
Zhijie Liu, Qiyi Tang, Sen Nie, Shi Wu, Liang Feng Zhang, Yutian Tang
摘要
Binary code similarity analysis (BCSA) is a crucial research area in many fields such as cybersecurity. Specifically, function-level diffing tools are the most widely used in BCSA: they perform function matching one by one for evaluating the similarity between binary programs. However, such methods need a high time complexity, making them unscalable in large-scale scenarios (e.g., 1/𝑛-to-𝑛 search). Towards effective and efficient program-level BCSA, we propose KEENHash, a novel hashing approach that hashes binaries into program-level representations through large language model (LLM)-generated function embeddings. KEENHash condenses a binary into one compact and fixed-length program embedding using K-Means and Feature Hashing, allowing us to do effective and efficient large-scale program-level BCSA, surpassing the previous state-of-the-art methods. The experimental results show that KEENHash is at least 215 times faster than the state-of-the-art function matching tools while maintaining effectiveness. Furthermore, in a large-scale scenario with 5.3 billion similarity evaluations, KEENHash takes only 395.83 seconds while these tools will cost at least 56 days. We also evaluate KEENHash on the program clone search of large-scale BCSA across extensive datasets in 202,305 binaries. Compared with 4 state-of-the-art methods, KEENHash outperforms all of them by at least 23.16%, and displays remarkable superiority over them in the large-scale BCSA security scenario of malware detection. CCS Concepts: • Security and privacy → Software reverse engineering.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Selective Knowledge Distillation: Fusing LLM Semantic Strengths with DNN Efficiency for Binary Code Similarity DetectionShize Zhou, Peiyu Liu, Lirong Fu, Tong Ye 等ACL 2026
- Towards Generality: Task-Adaptive Binary Analysis via Semantic Retrieval and Verifiable ReasoningYuzhe Liu, Zhijie Liu, Zhengmin Yu, Shu Wang 等USENIX Security 2026
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Understanding the Mirai BotnetManos Antonakakis, Tim April, Michael D. Bailey, Matt Bernhard 等USENIX Security 2017 · 被引用 2,003 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
相关 Paper
- CEBin: A Cost-Effective Framework for Large-Scale Binary Code Similarity DetectionHao Wang, Zeyu Gao, Chao Zhang, Mingyang Sun 等ISSTA 2024 · 被引用 21 次
- Transforming Generic Coder LLMs to Effective Binary Code Embedding Models for Similarity DetectionLitao Li, Leo Song, Steven H. H. Ding, Benjamin C. M. Fung 等NeurIPS 2025 · 被引用 2 次
- Scalable Program Clone Search through Spectral AnalysisTristan Benoit, Jean-Yves Marion, Sébastien BardinFSE 2023 · 被引用 5 次
- vSim: Semantics-Aware Value Extraction for Efficient Binary Code Similarity AnalysisHuaijin Wang, Zhiqiang LinNDSS 2026 · 被引用 3 次
- Neural Network-based Graph Embedding for Cross-Platform Binary Code Similarity DetectionXiaojun Xu, Chang Liu, Qian Feng, Heng Yin 等CCS 2017 · 被引用 682 次
