Lune

ISSTA2026Top-tier venue

Revisiting Graph Representations for ML-Based Binary Code Similarity Detection: A Systematic Study

Tengteng Yang, Yikun Hu, Jican Zhang, Lei Xue, Ming Fan, Liang Zhang

2026Year

Abstract

Binary Code Similarity Detection (BCSD) is a foundational capability in software security, underpinning critical applications ranging from vulnerability detection to malware analysis. While recent tools based on Machine Learning (ML) have achieved significant performance improvements, their efficacy is heavily contingent upon the underlying code representation. Through a systematic literature review of ML-based BCSD papers, we find that existing approaches typically leverage linear sequences or adopt graph-based representations, with the latter constituting the majority (77%). Despite this prevalence, there is no consensus on which graph representation yields superior effectiveness. Existing works usually couple graph construction with customized learning backbones and evaluate them on inconsistent benchmarks. This makes isolating the representation's impact difficult. Consequently, determining which graph topologies most effectively capture robust binary-code semantics under controlled and comparable evaluation settings remains an open problem. In this paper, we present a systematic study of graph representations for ML-based BCSD to bridge this gap. Specifically, we implement a modular evaluation framework that decouples graph construction from model training. Using this framework, we systematically evaluate seven representative graph representations, finding that no single representation is universally dominant, that distinct topologies exhibit unique strengths depending on the evaluation scenario, and that their rankings are largely backbone-stable despite varying absolute performance. We further investigate their combination effectiveness in N-day vulnerability detection and employ a tailored post-hoc analysis tool to study model-level structural reliance. The results show that DFG, PDG, and SOG subgraph pairs more often preserve trained models' similarity scores under pruning, while several other representations are more sensitive to structural reduction.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get bf7bc7cb-4de0-4cb4-ac24-c05203ec34ea

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines