"Len or index or count, anything but v1": Predicting Variable Names in Decompilation Output with Transfer Learning
Kuntal Kumar Pal, Ati Priya Bajaj, Pratyay Banerjee, Audrey Dutcher, Mutsumi Nakamura, Zion Leonahenahe Basque, Himanshu Gupta, Saurabh Arjun Sawant, Ujjwala Anantheswaran, Yan Shoshitaishvili, Adam Doupé, Chitta Baral, Ruoyu Wang
Abstract
Binary reverse engineering is an arduous and tedious task performed by skilled and expensive human analysts. Information about the source code is irrevocably lost in the compilation process. While modern decompilers attempt to generate C-style source code from a binary, they cannot recover lost variable names. Prior works have explored machine learning techniques for predicting variable names in decompiled code. However, the state-of-the-art systems, DIRE and DIRTY, generalize poorly to functions in the testing set that are not included in the training set—31.8% for DIRE on DIRTY’s data set and 36.9% for DIRTY on DIRTY’s data set.In this paper, we present VarBERT, a Bidirectional Encoder Representations from Transformers (BERT) to predict meaningful variable names in decompilation output. An advantage of VarBERT is that we can pre-train on human source code and then fine-tune the model to the task of predicting variable names. We also create a new data set VarCorpus, which significantly expands the size and variety of the data set. Our evaluation of VarBERT on VarCorpus, demonstrates a significant improvement in predicting the developer’s original variable names for O2 optimized binaries achieving accuracies of 54.43% for IDA and 54.49% for Ghidra. VarBERT is strictly better than state-of-the-art techniques: On a subset of VarCorpus, VarBERT could predict the developer’s original variable names 50.70% of the time, while DIRE and DIRTY predicted original variable names 35.94% and 38.00% of the time, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75d6361f-eec2-4cfd-88b1-6942eadc61fcCited by top-tier papers12
- Source Code Foundation Models are Transferable Binary Analysis Knowledge BasesZian Su, Xiangzhe Xu, Ziyang Huang, Kaiyuan Zhang et al.NeurIPS 2024 · 17 citations
- TYGR: Type Inference on Stripped Binaries using Graph Neural NetworksChang Zhu, Ziyang Li, Anton Xue, Ati Priya Bajaj et al.USENIX Security 2024 · 10 citations
- Idioms: A Simple and Effective Framework for Turbo-Charging Local Neural Decompilation with Well-Defined TypesLuke Dramko, Claire Le Goues, Edward J. SchwartzNDSS 2026 · 7 citations
- Bin2Wrong: a Unified Fuzzing Framework for Uncovering Semantic Errors in Binary-to-C DecompilersZao Yang, Stefan NagyUSENIX ATC 2025 · 6 citations
- FidelityGPT: Correcting Decompilation Distortions with Retrieval Augmented GenerationZhiping Zhou, Xiaohong Li, Ruitao Feng, Yao Zhang et al.NDSS 2026 · 5 citations
Builds on14
- SOK: (State of) The Art of War: Offensive Techniques in Binary AnalysisYan Shoshitaishvili, Ruoyu Wang, Christopher Salls, Nick Stephens et al.S&P 2016 · 1,085 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- Ramblr: Making Reassembly Great AgainRuoyu Wang, Yan Shoshitaishvili, Antonio Bianchi, Aravind Machiry et al.NDSS 2017 · 155 citations
- Debin: Predicting Debug Information in Stripped BinariesJingxuan He, Pesho Ivanov, Petar Tsankov, Veselin Raychev et al.CCS 2018 · 148 citations
- Helping Johnny to Analyze Malware: A Usability-Optimized Decompiler and Malware Analysis User StudyKhaled Yakdan, Sergej Dechand, Elmar Gerhards-Padilla, Matthew SmithS&P 2016 · 128 citations
Related papers
- Augmenting Decompiler Output with Learned Variable Names and TypesQibin Chen, Jeremy Lacomis, Edward J. Schwartz, Claire Le Goues et al.USENIX Security 2022
- RefBERT: A Two-Stage Pre-trained Framework for Automatic Rename RefactoringHao Liu, Yanlin Wang, Zhao Wei, Yong Xu et al.ISSTA 2023 · 19 citations
- DeGPT: Optimizing Decompiler Output with LLMPeiwei Hu, Ruigang Liang, Kai ChenNDSS 2024
- Unleashing the Power of Generative Model in Recovering Variable Names from Stripped BinaryXiangzhe Xu, Zhuo Zhang, Zian Su, Ziyang Huang et al.NDSS 2025
- HexT5: Unified Pre-Training for Stripped Binary Code Information InferenceJiaqi Xiong, Guoqiang Chen, Kejiang Chen, Han Gao et al.ASE 2023 · 8 citations
