Unveiling the Fragility of Binary Code Similarity Detection via Targeted Attacks with Model Explanations
Mingjie Chen, Tiancheng Zhu, Mingxue Zhang, Yiling He, Minghao Lin, Penghui Li, Kui Ren
摘要
Binary code similarity detection (BCSD) serves as a fundamental technique for various software engineering tasks, e.g., vulnerability detection and classification. Attacks against such BCSD models have therefore drawn extensive attention, aiming at misleading the models to generate erroneous predictions. Prior works have explored various approaches to generating semantic-preserving variants, i.e., adversarial samples, to evaluate the robustness of the models against adversarial attacks. However, they have mainly relied on heuristic criteria or iterative greedy algorithms to locate salient code influencing the model output, which often leads to inefficient search and high computational cost. Moreover, when processing programs with high complexities, such attacks tend to be time-consuming.
In this work, we unveil the fragility of BCSD models through a novel attack framework guided by model explanations. In particular, we focus on targeted attacks where the attack goal is to mislead the model's predictions to a specific target. Our attack leverages explainers to pinpoint critical code snippet for perturbations, reducing the exploration overhead. The evaluation results demonstrate that the proposed attacks effectively improve the attack efficiency, while maintaining comparable or higher success rates. Importantly, the speedup for perturbation target selection achieves up to 63.66×, demonstrating the practical value of explanation-guided localization. Our real-world case studies on vulnerability detection and classification further demonstrate the security implications of our attacks, highlighting fundamental robustness limitations in current BCSD models, and the urgent need for more robust designs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- Parameterized Explainer for Graph Neural NetworkDongsheng Luo, Wei Cheng, Dongkuan Xu, Wenchao Yu 等NeurIPS 2020 · 被引用 888 次
- Neural Network-based Graph Embedding for Cross-Platform Binary Code Similarity DetectionXiaojun Xu, Chang Liu, Qian Feng, Heng Yin 等CCS 2017 · 被引用 682 次
- LEMNA: Explaining Deep Learning based Security ApplicationsWenbo Guo, Dongliang Mu, Jun Xu, Purui Su 等CCS 2018 · 被引用 336 次
- Intriguing Properties of Adversarial ML Attacks in the Problem SpaceFabio Pierazzi, Feargus Pendlebury, Jacopo Cortellazzi, Lorenzo CavallaroS&P 2020 · 被引用 334 次
相关 Paper
- CEBin: A Cost-Effective Framework for Large-Scale Binary Code Similarity DetectionHao Wang, Zeyu Gao, Chao Zhang, Mingyang Sun 等ISSTA 2024 · 被引用 21 次
- Fool Me If You Can: On the Robustness of Binary Code Similarity Detection Models against Semantics-Preserving TransformationsJiyong Uhm, Minseok Kim, Michalis Polychronakis, Hyungjoon KooFSE 2026 · 被引用 1 次
- Advancing Binary Code Similarity Detection via Context-Content Fusion and LLM VerificationChaopeng Dong, Jingdong Guo, Shouguo Yang, Yi Li 等ASE 2025
- Revisiting Graph Representations for ML-Based Binary Code Similarity Detection: A Systematic StudyTengteng Yang, Yikun Hu, Jican Zhang, Lei Xue 等ISSTA 2026
- Improving ML-based Binary Function Similarity Detection by Assessing and Deprioritizing Control Flow Graph FeaturesJialai Wang, Chao Zhang, Longfei Chen, Yi Rong 等USENIX Security 2024 · 被引用 15 次
