Can a Deep Learning Model for One Architecture Be Used for Others? Retargeted-Architecture Binary Code Analysis
Junzhe Wang, Matthew Sharp, Chuxiong Wu, Qiang Zeng, Lannan Luo
摘要
NLP-inspired deep learning for binary code analysis demonstrates notable performance. Considering the diverse Instruction Set Architectures (ISAs) on the market, it is important to be able to analyze code of various ISAs. However, training a deep learning model usually requires a large amount of data, which poses a challenge for certain ISAs such as PowerPC that suffer from the "data scarcity" issue. For instance, acquiring a large dataset of PowerPC malware proves to be challenging. Moreover, given a binary analysis task and multiple ISAs, it takes much time and effort (e.g., for data collection, labeling and cleaning, and parameter tuning) to train one model per ISA. We propose a new direction, retargeted-architecture binary code analysis, to handle the data scarcity issue and alleviate the per-ISA effort. Our idea is to transfer knowledge from one ISA to others—that is, a model, trained with rich data and much time and effort for one ISA, can perform prediction for others without any modification. We showcase the idea through two important tasks: malware detection and function similarity detection. An extensive evaluation involving four ISAs (x86, ARM, MIPS, and PowerPC) demonstrates the effectiveness of the approach and the high performance is interpreted.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- From One Thousand Pages of Specification to Unveiling Hidden Bugs: Large Language Model Assisted Fuzzing of Matter IoT DevicesXiaoyue Ma, Lannan Luo, Qiang ZengUSENIX Security 2024 · 被引用 49 次
- ReSym: Harnessing LLMs to Recover Variable and Data Structure Symbols from Stripped BinariesDanning Xie, Zhuo Zhang, Nan Jiang, Xiangzhe Xu 等CCS 2024 · 被引用 21 次
- Beyond Raw Bytes: Towards Large Malware Language ModelsLuke Kurlandski, Harel Berger, Yin Pan, Matthew WrightNDSS 2026 · 被引用 5 次
- A Large Scale Study of AI-based Binary Function Similarity Detection Techniques for Security Researchers and PractitionersJingyi Shi, Yufeng Chen, Yang Xiao, Yuekang Li 等ASE 2025
- You Can't Judge a Binary by Its Header: Data-Code Separation for Non-Standard ARM Binaries Using Pseudo LabelsHadjer Benkraouda, Nirav Diwan, Gang WangS&P 2025
它引用的顶会 Paper11
- Neural Network-based Graph Embedding for Cross-Platform Binary Code Similarity DetectionXiaojun Xu, Chang Liu, Qian Feng, Heng Yin 等CCS 2017 · 被引用 682 次
- Scalable Graph-based Bug Search for Firmware ImagesQian Feng, Rundong Zhou, Chengcheng Xu, Yao Cheng 等CCS 2016 · 被引用 456 次
- Asm2Vec: Boosting Static Representation Robustness for Binary Clone Search against Code Obfuscation and Compiler OptimizationSteven H. H. Ding, Benjamin C. M. Fung, Philippe CharlandS&P 2019 · 被引用 447 次
- discovRE: Efficient Cross-Architecture Identification of Bugs in Binary CodeSebastian Eschweiler, Khaled Yakdan, Elmar Gerhards-PadillaNDSS 2016 · 被引用 342 次
- Neural Machine Translation Inspired Binary Code Similarity Comparison beyond Function PairsFei Zuo, Xiaopeng Li, Patrick Young, Lannan Luo 等NDSS 2019 · 被引用 262 次
相关 Paper
- PalmTree: Learning an Assembly Language Model for Instruction EmbeddingXuezixiang Li, Yu Qu, Heng YinCCS 2021 · 被引用 139 次
- Improving Binary Code Similarity Transformer Models by Semantics-Driven Instruction DeemphasisXiangzhe Xu, Shiwei Feng, Yapeng Ye, Guangyu Shen 等ISSTA 2023 · 被引用 25 次
- Fool Me If You Can: On the Robustness of Binary Code Similarity Detection Models against Semantics-Preserving TransformationsJiyong Uhm, Minseok Kim, Michalis Polychronakis, Hyungjoon KooFSE 2026 · 被引用 1 次
- BinAug: Enhancing Binary Similarity Analysis with Low-Cost Input RepairingWai Kin Wong, Huaijin Wang, Zongjie Li, Shuai WangICSE 2024 · 被引用 4 次
- DEEPVSA: Facilitating Value-set Analysis with Deep Learning for Postmortem Program AnalysisWenbo Guo, Dongliang Mu, Xinyu Xing, Min Du 等USENIX Security 2019 · 被引用 67 次
