Unveiling the Fragility of Binary Code Similarity Detection via Targeted Attacks with Model Explanations
Mingjie Chen, Tiancheng Zhu, Mingxue Zhang, Yiling He, Minghao Lin, Penghui Li, Kui Ren
Abstract
Binary code similarity detection (BCSD) serves as a fundamental technique for various software engineering tasks, e.g., vulnerability detection and classification. Attacks against such BCSD models have therefore drawn extensive attention, aiming at misleading the models to generate erroneous predictions. Prior works have explored various approaches to generating semantic-preserving variants, i.e., adversarial samples, to evaluate the robustness of the models against adversarial attacks. However, they have mainly relied on heuristic criteria or iterative greedy algorithms to locate salient code influencing the model output, which often leads to inefficient search and high computational cost. Moreover, when processing programs with high complexities, such attacks tend to be time-consuming.
In this work, we unveil the fragility of BCSD models through a novel attack framework guided by model explanations. In particular, we focus on targeted attacks where the attack goal is to mislead the model's predictions to a specific target. Our attack leverages explainers to pinpoint critical code snippet for perturbations, reducing the exploration overhead. The evaluation results demonstrate that the proposed attacks effectively improve the attack efficiency, while maintaining comparable or higher success rates. Importantly, the speedup for perturbation target selection achieves up to 63.66×, demonstrating the practical value of explanation-guided localization. Our real-world case studies on vulnerability detection and classification further demonstrate the security implications of our attacks, highlighting fundamental robustness limitations in current BCSD models, and the urgent need for more robust designs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9fe7e895-5ccb-49ee-b52b-e99af06760d1Builds on16
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Parameterized Explainer for Graph Neural NetworkDongsheng Luo, Wei Cheng, Dongkuan Xu, Wenchao Yu et al.NeurIPS 2020 · 888 citations
- Neural Network-based Graph Embedding for Cross-Platform Binary Code Similarity DetectionXiaojun Xu, Chang Liu, Qian Feng, Heng Yin et al.CCS 2017 · 682 citations
- LEMNA: Explaining Deep Learning based Security ApplicationsWenbo Guo, Dongliang Mu, Jun Xu, Purui Su et al.CCS 2018 · 336 citations
- Intriguing Properties of Adversarial ML Attacks in the Problem SpaceFabio Pierazzi, Feargus Pendlebury, Jacopo Cortellazzi, Lorenzo CavallaroS&P 2020 · 334 citations
Related papers
- CEBin: A Cost-Effective Framework for Large-Scale Binary Code Similarity DetectionHao Wang, Zeyu Gao, Chao Zhang, Mingyang Sun et al.ISSTA 2024 · 21 citations
- Fool Me If You Can: On the Robustness of Binary Code Similarity Detection Models against Semantics-Preserving TransformationsJiyong Uhm, Minseok Kim, Michalis Polychronakis, Hyungjoon KooFSE 2026 · 1 citation
- Advancing Binary Code Similarity Detection via Context-Content Fusion and LLM VerificationChaopeng Dong, Jingdong Guo, Shouguo Yang, Yi Li et al.ASE 2025
- Revisiting Graph Representations for ML-Based Binary Code Similarity Detection: A Systematic StudyTengteng Yang, Yikun Hu, Jican Zhang, Lei Xue et al.ISSTA 2026
- Improving ML-based Binary Function Similarity Detection by Assessing and Deprioritizing Control Flow Graph FeaturesJialai Wang, Chao Zhang, Longfei Chen, Yi Rong et al.USENIX Security 2024 · 15 citations
