Neural reverse engineering of stripped binaries using augmented control flow graphs
Yaniv David, Uri Alon, Eran Yahav
Abstract
We address the problem of reverse engineering of stripped executables, which contain no debug information. This is a challenging problem because of the low amount of syntactic information available in stripped executables, and the diverse assembly code patterns arising from compiler optimizations.
We present a novel approach for predicting procedure names in stripped executables. Our approach combines static analysis with neural models. The main idea is to use static analysis to obtain augmented representations of call sites; encode the structure of these call sites using the control-flow graph (CFG) and finally, generate a target name while attending to these call sites. We use our representation to drive graph-based, LSTM-based and Transformer-based architectures.
Our evaluation shows that our models produce predictions that are difficult and time consuming for humans, while improving on existing methods by 28% and by 100% over state-of-the-art neural textual models that do not use any static analysis. Code and data for this evaluation are available at https://github.com/tech-srl/Nero.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da7d1b28-983f-422d-853f-0bfb5b3d9227Cited by top-tier papers32
- DOBF: A Deobfuscation Pre-Training Objective for Programming LanguagesMarie-Anne Lachaux, Baptiste Rozière, Marc Szafraniec, Guillaume LampleNeurIPS 2021 · 174 citations
- Adversarial examples for models of codeNoam Yefet, Uri Alon, Eran YahavOOPSLA 2020 · 162 citations
- Typilus: neural type hintsMiltiadis Allamanis, Earl T. Barr, Soline Ducousso, Zheng GaoPLDI 2020 · 92 citations
- How could Neural Networks understand Programs?Dinglan Peng, Shuxin Zheng, Yatao Li, Guolin Ke et al.ICML 2021 · 75 citations
- SymLM: Predicting Function Names in Stripped Binaries via Context-Sensitive Execution-Aware Code EmbeddingsXin Jin, Kexin Pei, Jun Yeon Won, Zhiqiang LinCCS 2022 · 56 citations
Builds on5
- Neural Network-based Graph Embedding for Cross-Platform Binary Code Similarity DetectionXiaojun Xu, Chang Liu, Qian Feng, Heng Yin et al.CCS 2017 · 682 citations
- Asm2Vec: Boosting Static Representation Robustness for Binary Clone Search against Code Obfuscation and Compiler OptimizationSteven H. H. Ding, Benjamin C. M. Fung, Philippe CharlandS&P 2019 · 447 citations
- Debin: Predicting Debug Information in Stripped BinariesJingxuan He, Pesho Ivanov, Petar Tsankov, Veselin Raychev et al.CCS 2018 · 148 citations
- Structural Language Models of CodeUri Alon, Roy Sadaka, Omer Levy, Eran YahavICML 2020 · 115 citations
- An Observational Investigation of Reverse Engineers' ProcessesDaniel Votipka, Seth M. Rabin, Kristopher K. Micinski, Jeffrey S. Foster et al.USENIX Security 2020
Related papers
- A lightweight framework for function name reassignment based on large-scale stripped binariesHan Gao, Shaoyin Cheng, Yinxing Xue, Weiming ZhangISSTA 2021 · 46 citations
- BinRAG: An RAG-Based Decompilation Framework Fusing Name Prediction and Calling ContextWai Kin Wong, Daoyuan Wu, Zhibo Liu, Huaijin Wang et al.ISSTA 2026
- Hieronym: Leveraging Hierarchical Multi-Source Information for Function Renaming in Stripped BinaryXiaoling Zhang, Jian Sun, Dawei Wang, Chongyu Wang et al.CCS 2026
- Accurate Disassembly of Complex Binaries Without Use of Compiler MetadataSoumyakant Priyadarshan, Huan Nguyen, R. SekarASPLOS 2023 · 7 citations
- TYGR: Type Inference on Stripped Binaries using Graph Neural NetworksChang Zhu, Ziyang Li, Anton Xue, Ati Priya Bajaj et al.USENIX Security 2024 · 10 citations
