Advancing Binary Code Similarity Detection via Context-Content Fusion and LLM Verification
Chaopeng Dong, Jingdong Guo, Shouguo Yang, Yi Li, Dongliang Fang, Yang Xiao, Yongle Chen, Limin Sun
摘要
Binary Code Similarity Detection (BCSD), essential for binary-code related tasks like vulnerability detection, has attracted increasing attention in recent years. However, existing methods frequently fall short of achieving both high precision and recall at scale, and their results often lack interpretability due to the neglect of function context and reliance on purely similarity-driven outputs. Our key insights are twofold: 1) Binary functions are not self-contained; they depend on other code and data beyond their content to fulfill their functionalities. 2) Large language models (LLMs) excel not only at analyzing code but also at generating reasonable explanations. Motivated by these insights, we propose a general BCSD framework, Co<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup>F uLL. We first systematically select stable and representative code and data features, along with their corresponding dependencies on the functions, to construct the function context. Then, by fusing function context with content similarities computed by the existing BCSD approach, we substantially narrow down the search space. Ultimately, we employ LLMs with a carefully designed prompt to verify the remaining candidates and produce clear, human-readable explanations. We conduct comprehensive experiments on a large function pool under varying compilation settings and after binary stripping. The results show that Co2F uLL based on HermesSim and DeepSeek-V3 achieves 80.5% precision and 94.4% recall, improving the baseline HermesSim by 142.5% and 42.2%, respectively, providing an accurate and interpretable solution for BCSD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Neural Network-based Graph Embedding for Cross-Platform Binary Code Similarity DetectionXiaojun Xu, Chang Liu, Qian Feng, Heng Yin 等CCS 2017 · 被引用 682 次
- Language Models can Solve Computer TasksGeunwoo Kim, Pierre Baldi, Stephen McAleerNeurIPS 2023 · 被引用 539 次
- jTrans: jump-aware transformer for binary code similarity detectionHao Wang, Wenjie Qu, Gilad Katz, Wenyu Zhu 等ISSTA 2022 · 被引用 139 次
相关 Paper
- Selective Knowledge Distillation: Fusing LLM Semantic Strengths with DNN Efficiency for Binary Code Similarity DetectionShize Zhou, Peiyu Liu, Lirong Fu, Tong Ye 等ACL 2026
- Enhancing Vulnerability Detection via Inter-procedural Semantic CompletionBozhi Wu, Chengjie Liu, Zhiming Li, Yushi Cao 等ISSTA 2025 · 被引用 2 次
- CEBin: A Cost-Effective Framework for Large-Scale Binary Code Similarity DetectionHao Wang, Zeyu Gao, Chao Zhang, Mingyang Sun 等ISSTA 2024 · 被引用 21 次
- Transforming Generic Coder LLMs to Effective Binary Code Embedding Models for Similarity DetectionLitao Li, Leo Song, Steven H. H. Ding, Benjamin C. M. Fung 等NeurIPS 2025 · 被引用 2 次
- Code is not Natural Language: Unlock the Power of Semantics-Oriented Graph Representation for Binary Code Similarity DetectionHaojie He, Xingwei Lin, Ziang Weng, Ruijie Zhao 等USENIX Security 2024 · 被引用 66 次
