"Len or index or count, anything but v1": Predicting Variable Names in Decompilation Output with Transfer Learning
Kuntal Kumar Pal, Ati Priya Bajaj, Pratyay Banerjee, Audrey Dutcher, Mutsumi Nakamura, Zion Leonahenahe Basque, Himanshu Gupta, Saurabh Arjun Sawant, Ujjwala Anantheswaran, Yan Shoshitaishvili, Adam Doupé, Chitta Baral, Ruoyu Wang
摘要
Binary reverse engineering is an arduous and tedious task performed by skilled and expensive human analysts. Information about the source code is irrevocably lost in the compilation process. While modern decompilers attempt to generate C-style source code from a binary, they cannot recover lost variable names. Prior works have explored machine learning techniques for predicting variable names in decompiled code. However, the state-of-the-art systems, DIRE and DIRTY, generalize poorly to functions in the testing set that are not included in the training set—31.8% for DIRE on DIRTY’s data set and 36.9% for DIRTY on DIRTY’s data set.In this paper, we present VarBERT, a Bidirectional Encoder Representations from Transformers (BERT) to predict meaningful variable names in decompilation output. An advantage of VarBERT is that we can pre-train on human source code and then fine-tune the model to the task of predicting variable names. We also create a new data set VarCorpus, which significantly expands the size and variety of the data set. Our evaluation of VarBERT on VarCorpus, demonstrates a significant improvement in predicting the developer’s original variable names for O2 optimized binaries achieving accuracies of 54.43% for IDA and 54.49% for Ghidra. VarBERT is strictly better than state-of-the-art techniques: On a subset of VarCorpus, VarBERT could predict the developer’s original variable names 50.70% of the time, while DIRE and DIRTY predicted original variable names 35.94% and 38.00% of the time, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Source Code Foundation Models are Transferable Binary Analysis Knowledge BasesZian Su, Xiangzhe Xu, Ziyang Huang, Kaiyuan Zhang 等NeurIPS 2024 · 被引用 17 次
- TYGR: Type Inference on Stripped Binaries using Graph Neural NetworksChang Zhu, Ziyang Li, Anton Xue, Ati Priya Bajaj 等USENIX Security 2024 · 被引用 10 次
- Idioms: A Simple and Effective Framework for Turbo-Charging Local Neural Decompilation with Well-Defined TypesLuke Dramko, Claire Le Goues, Edward J. SchwartzNDSS 2026 · 被引用 7 次
- Bin2Wrong: a Unified Fuzzing Framework for Uncovering Semantic Errors in Binary-to-C DecompilersZao Yang, Stefan NagyUSENIX ATC 2025 · 被引用 6 次
- FidelityGPT: Correcting Decompilation Distortions with Retrieval Augmented GenerationZhiping Zhou, Xiaohong Li, Ruitao Feng, Yao Zhang 等NDSS 2026 · 被引用 5 次
它引用的顶会 Paper14
- SOK: (State of) The Art of War: Offensive Techniques in Binary AnalysisYan Shoshitaishvili, Ruoyu Wang, Christopher Salls, Nick Stephens 等S&P 2016 · 被引用 1,085 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- Ramblr: Making Reassembly Great AgainRuoyu Wang, Yan Shoshitaishvili, Antonio Bianchi, Aravind Machiry 等NDSS 2017 · 被引用 155 次
- Debin: Predicting Debug Information in Stripped BinariesJingxuan He, Pesho Ivanov, Petar Tsankov, Veselin Raychev 等CCS 2018 · 被引用 148 次
- Helping Johnny to Analyze Malware: A Usability-Optimized Decompiler and Malware Analysis User StudyKhaled Yakdan, Sergej Dechand, Elmar Gerhards-Padilla, Matthew SmithS&P 2016 · 被引用 128 次
相关 Paper
- Augmenting Decompiler Output with Learned Variable Names and TypesQibin Chen, Jeremy Lacomis, Edward J. Schwartz, Claire Le Goues 等USENIX Security 2022
- RefBERT: A Two-Stage Pre-trained Framework for Automatic Rename RefactoringHao Liu, Yanlin Wang, Zhao Wei, Yong Xu 等ISSTA 2023 · 被引用 19 次
- DeGPT: Optimizing Decompiler Output with LLMPeiwei Hu, Ruigang Liang, Kai ChenNDSS 2024
- Unleashing the Power of Generative Model in Recovering Variable Names from Stripped BinaryXiangzhe Xu, Zhuo Zhang, Zian Su, Ziyang Huang 等NDSS 2025
- HexT5: Unified Pre-Training for Stripped Binary Code Information InferenceJiaqi Xiong, Guoqiang Chen, Kejiang Chen, Han Gao 等ASE 2023 · 被引用 8 次
