Lune

NeurIPS2025顶会

Transforming Generic Coder LLMs to Effective Binary Code Embedding Models for Similarity Detection

Litao Li, Leo Song, Steven H. H. Ding, Benjamin C. M. Fung, Philippe Charland

2025年份
2被引次数

摘要

Cybersecurity and software research have crossed paths with modern deep learning research for a few years. The power of large language models (LLMs) in particular has intrigued us to apply them to understanding binary code. In this paper, we investigate some of the many ways LLMs can be applied to binary code similarity detection, as it is a significantly more difficult task compared to source code similarity detection due to the sparsity of information and less meaningful syntax. It also has great practical implications, such as vulnerability and malware detection. We find that pretrained LLMs are mostly capable of detecting similar binary code, even with a zero-shot setting. Our main contributions and findings are to provide several supervised fine-tuning methods that, when combined, significantly surpass zero-shot LLMs and state-of-the-art binary code similarity detection methods. Specifically, we up-train the model through data augmentation, translation-style causal learning, LLM2Vec, and cumulative GTE loss. With a complete ablation study, we show that our training method can transform a generic language model into a powerful binary similarity expert, and is also robust and general enough for cross-optimization, cross-architecture, and cross-obfuscation detection.

Many binary code models are trained from scratch [Yang et al., 2021;Tian et al., 2021;Yu et al., 2020] and their sizes are much smaller compared to LLMs. They typically focus on cross-optimization retrieval and fail to generalize to diverse compiler settings. Other works take pre-trained LLMs [Tan et al., 2024;Wang et al., 2022Wang et al., , 2024] ] and apply custom fine-tuning techniques for binary code matching. However, most of these approaches rely on either closed-source models like GPT [Radford, 2018] or large model sizes. This leads to scalability issues when the computational resource is a constraint, which realistically is the case with small research labs or even companies. We want to explore effective and efficient LLM fine-tuning for binary code embedding and matching.

In this work, we address the aforementioned problems and propose EBM (Effective Binary Matching), our novel training framework, including carefully chosen data augmentation and fine-tuning processes, to uptrain a generic LLM into a binary code embedding and matching expert. We show in our experiments that LLMs have become dominant enough that even zero-shot models can surpass welltrained binary code models. With our fine-tuning applied, EBM can significantly increase the mean reciprocal rank (MRR) by 10% to 70%, depending on the tasks. We also provide a comprehensive ablation study to prove and emphasize the importance of each training process. The dataset and code can be accessed on Github 1 . Our major contributions are:

• We propose a multi-training framework to utilize the power of generic LLMs specifically for BCSD. We target different compiler settings, including cross-optimization, crossarchitecture, and cross-obfuscation, to build a general and effective similarity retriever.

• To combat the lack of assembly code LLMs, we uptrain the generic LLM to a binary code-specific model. Particularly in cross-architecture similarity detection, this allows for significantly better translation between different syntaxes.

• We utilize LLM2Vec to build a refinement of the assembly tokens, which encodes better semantics through a masked next token prediction task.

• In the downstream contrastive learning task, we propose an enhanced version of InfoNCE loss to utilize all available samples within the batch. This is particularly useful when computing resources are limited and large models are trained.

• We build and compare various baselines to evaluate the effect of different language/coder models and state-of-the-art BCSD models. In our similarity retrieval evaluation, our approach outperforms all benchmarks for both datasets.

• We conduct thorough ablation studies and in-depth analysis to investigate all our training tasks and how they contribute to the similarity retrieval result. We show that all our training tasks are essential and can improve retrieval performance.

Traditional BCSD Without data-driven or learning-based methods, code similarity detection is traditionally conducted using static analysis, dynamic analysis, or code-based algorithms. Static analysis usually involves graph matching [Dullien and Rolles, 2005;Bourquin et al., 2013], where control flow graphs are extracted from assembly code and compared using algorithms or user-defined heuristics. Dynamic analysis instead leverages runtime or symbolic execution to investigate program behavior [

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper14

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖