Transforming Generic Coder LLMs to Effective Binary Code Embedding Models for Similarity Detection
Litao Li, Leo Song, Steven H. H. Ding, Benjamin C. M. Fung, Philippe Charland
Abstract
Cybersecurity and software research have crossed paths with modern deep learning research for a few years. The power of large language models (LLMs) in particular has intrigued us to apply them to understanding binary code. In this paper, we investigate some of the many ways LLMs can be applied to binary code similarity detection, as it is a significantly more difficult task compared to source code similarity detection due to the sparsity of information and less meaningful syntax. It also has great practical implications, such as vulnerability and malware detection. We find that pretrained LLMs are mostly capable of detecting similar binary code, even with a zero-shot setting. Our main contributions and findings are to provide several supervised fine-tuning methods that, when combined, significantly surpass zero-shot LLMs and state-of-the-art binary code similarity detection methods. Specifically, we up-train the model through data augmentation, translation-style causal learning, LLM2Vec, and cumulative GTE loss. With a complete ablation study, we show that our training method can transform a generic language model into a powerful binary similarity expert, and is also robust and general enough for cross-optimization, cross-architecture, and cross-obfuscation detection.
Many binary code models are trained from scratch [Yang et al., 2021;Tian et al., 2021;Yu et al., 2020] and their sizes are much smaller compared to LLMs. They typically focus on cross-optimization retrieval and fail to generalize to diverse compiler settings. Other works take pre-trained LLMs [Tan et al., 2024;Wang et al., 2022Wang et al., , 2024] ] and apply custom fine-tuning techniques for binary code matching. However, most of these approaches rely on either closed-source models like GPT [Radford, 2018] or large model sizes. This leads to scalability issues when the computational resource is a constraint, which realistically is the case with small research labs or even companies. We want to explore effective and efficient LLM fine-tuning for binary code embedding and matching.
In this work, we address the aforementioned problems and propose EBM (Effective Binary Matching), our novel training framework, including carefully chosen data augmentation and fine-tuning processes, to uptrain a generic LLM into a binary code embedding and matching expert. We show in our experiments that LLMs have become dominant enough that even zero-shot models can surpass welltrained binary code models. With our fine-tuning applied, EBM can significantly increase the mean reciprocal rank (MRR) by 10% to 70%, depending on the tasks. We also provide a comprehensive ablation study to prove and emphasize the importance of each training process. The dataset and code can be accessed on Github 1 . Our major contributions are:
• We propose a multi-training framework to utilize the power of generic LLMs specifically for BCSD. We target different compiler settings, including cross-optimization, crossarchitecture, and cross-obfuscation, to build a general and effective similarity retriever.
• To combat the lack of assembly code LLMs, we uptrain the generic LLM to a binary code-specific model. Particularly in cross-architecture similarity detection, this allows for significantly better translation between different syntaxes.
• We utilize LLM2Vec to build a refinement of the assembly tokens, which encodes better semantics through a masked next token prediction task.
• In the downstream contrastive learning task, we propose an enhanced version of InfoNCE loss to utilize all available samples within the batch. This is particularly useful when computing resources are limited and large models are trained.
• We build and compare various baselines to evaluate the effect of different language/coder models and state-of-the-art BCSD models. In our similarity retrieval evaluation, our approach outperforms all benchmarks for both datasets.
• We conduct thorough ablation studies and in-depth analysis to investigate all our training tasks and how they contribute to the similarity retrieval result. We show that all our training tasks are essential and can improve retrieval performance.
Traditional BCSD Without data-driven or learning-based methods, code similarity detection is traditionally conducted using static analysis, dynamic analysis, or code-based algorithms. Static analysis usually involves graph matching [Dullien and Rolles, 2005;Bourquin et al., 2013], where control flow graphs are extracted from assembly code and compared using algorithms or user-defined heuristics. Dynamic analysis instead leverages runtime or symbolic execution to investigate program behavior [
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f35a7177-91f3-4985-97ec-5a4e26ea3a59Builds on14
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Neural Network-based Graph Embedding for Cross-Platform Binary Code Similarity DetectionXiaojun Xu, Chang Liu, Qian Feng, Heng Yin et al.CCS 2017 · 682 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
Related papers
- Selective Knowledge Distillation: Fusing LLM Semantic Strengths with DNN Efficiency for Binary Code Similarity DetectionShize Zhou, Peiyu Liu, Lirong Fu, Tong Ye et al.ACL 2026
- Improving Binary Code Similarity Transformer Models by Semantics-Driven Instruction DeemphasisXiangzhe Xu, Shiwei Feng, Yapeng Ye, Guangyu Shen et al.ISSTA 2023 · 25 citations
- Nova: Generative Language Models for Assembly Code with Hierarchical Attention and Contrastive LearningNan Jiang, Chengxiao Wang, Kevin Liu, Xiangzhe Xu et al.ICLR 2025
- Advancing Binary Code Similarity Detection via Context-Content Fusion and LLM VerificationChaopeng Dong, Jingdong Guo, Shouguo Yang, Yi Li et al.ASE 2025
- Beyond Classification: Inferring Function Names in Stripped Binaries via Domain Adapted LLMsLinxi Jiang, Xin Jin, Zhiqiang LinNDSS 2025
