StateFormer: fine-grained type recovery from binaries using generative state modeling
Kexin Pei, Jonas Guan, Matthew Broughton, Zhongtian Chen, Songchen Yao, David Williams-King, Vikas Ummadisetty, Junfeng Yang, Baishakhi Ray, Suman Jana
Abstract
Binary type inference is a critical reverse engineering task supporting many security applications, including vulnerability analysis, binary hardening, forensics, and decompilation. It is a difficult task because source-level type information is often stripped during compilation, leaving only binaries with untyped memory and register accesses. Existing approaches rely on hand-coded type inference rules defined by domain experts, which are brittle and require nontrivial effort to maintain and update. Even though machine learning approaches have shown promise at automatically learning the inference rules, their accuracy is still low, especially for optimized binaries.
We present STATEFORMER, a new neural architecture that is adept at accurate and robust type inference. STATEFORMER follows a twostep transfer learning paradigm. In the pretraining step, the model is trained with Generative State Modeling (GSM), a novel task that we design to teach the model to statically approximate execution effects of assembly instructions in both forward and backward directions. In the finetuning step, the pretrained model learns to use its knowledge of operational semantics to infer types.
We evaluate STATEFORMER's performance on a corpus of 33 popular open-source software projects containing over 1.67 billion variables of different types. The programs are compiled with GCC and LLVM over 4 optimization levels O0-O3, and 3 obfuscation passes based on LLVM. Our model significantly outperforms stateof-the-art ML-based tools by 14.6% in recovering types for both function arguments and variables. Our ablation studies show that GSM improves type inference accuracy by 33%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b1aee71-5c17-45fa-a7d0-956eb4ac7862Cited by top-tier papers35
- SymLM: Predicting Function Names in Stripped Binaries via Context-Sensitive Execution-Aware Code EmbeddingsXin Jin, Kexin Pei, Jun Yeon Won, Zhiqiang LinCCS 2022 · 56 citations
- Finding the dwarf: recovering precise types from WebAssembly binariesDaniel Lehmann, Michael PradelPLDI 2022 · 29 citations
- TRACED: Execution-aware Pre-training for Source CodeYangruibo Ding, Benjamin Steenhoek, Kexin Pei, Gail E. Kaiser et al.ICSE 2024 · 29 citations
- DeepInfer: Deep Type Inference from Smart Contract BytecodeKunsong Zhao, Zihao Li, Jianfeng Li, He Ye et al.FSE 2023 · 24 citations
- A Taxonomy of C Decompiler Fidelity IssuesLuke Dramko, Jeremy Lacomis, Edward J. Schwartz, Bogdan Vasilescu et al.USENIX Security 2024 · 23 citations
Builds on18
- SOK: (State of) The Art of War: Offensive Techniques in Binary AnalysisYan Shoshitaishvili, Ruoyu Wang, Christopher Salls, Nick Stephens et al.S&P 2016 · 1,085 citations
- A Tough Call: Mitigating Advanced Code-Reuse Attacks at the Binary LevelVictor van der Veen, Enes Göktas, Moritz Contag, Andre Pawlowski et al.S&P 2016 · 227 citations
- Neural Nets Can Learn Function Type Signatures From BinariesZheng Leong Chua, Shiqi Shen, Prateek Saxena, Zhenkai LiangUSENIX Security 2017 · 175 citations
- Debin: Predicting Debug Information in Stripped BinariesJingxuan He, Pesho Ivanov, Petar Tsankov, Veselin Raychev et al.CCS 2018 · 148 citations
- Where Does It Go?: Refining Indirect-Call Targets with Multi-Layer Type AnalysisKangjie Lu, Hong HuCCS 2019 · 142 citations
Related papers
- TYGR: Type Inference on Stripped Binaries using Graph Neural NetworksChang Zhu, Ziyang Li, Anton Xue, Ati Priya Bajaj et al.USENIX Security 2024 · 10 citations
- TypeForge: Synthesizing and Selecting Best-Fit Composite Data Types for Stripped BinariesYanzhong Wang, Ruigang Liang, Yilin Li, Peiwei Hu et al.S&P 2025
- Beyond Classification: Inferring Function Names in Stripped Binaries via Domain Adapted LLMsLinxi Jiang, Xin Jin, Zhiqiang LinNDSS 2025
- NeuDep: neural binary memory dependence analysisKexin Pei, Dongdong She, Michael Wang, Scott Geng et al.FSE 2022 · 8 citations
- Hieronym: Leveraging Hierarchical Multi-Source Information for Function Renaming in Stripped BinaryXiaoling Zhang, Jian Sun, Dawei Wang, Chongyu Wang et al.CCS 2026
