Idioms: A Simple and Effective Framework for Turbo-Charging Local Neural Decompilation with Well-Defined Types
Luke Dramko, Claire Le Goues, Edward J. Schwartz
Abstract
Decompilers help reverse engineers analyze software at a higher level of abstraction than assembly code. Unfortunately, because compilation is lossy, traditional decompilers, which are deterministic, produce code that lacks many characteristics that make source code readable in the first place, such as variable and type names. Neural decompilers offer the exciting possibility of statistically filling in these details. Unfortunately, existing work in neural decompilation suffers from substantial limitations that preclude its use on real code, such as the inability to provide definitions for user-defined composite types. In this work, we introduce Idioms, a simple, generalizable, and effective neural decompilation approach that can finetune any LLM into a neural decompiler capable of generating the appropriate user-defined type definitions alongside the decompiled code, and a new dataset, Realtype, that includes substantially more complicated and realistic types than existing neural decompilation benchmarks. We show that our approach yields state-of-the-art results in neural decompilation. On the most challenging existing benchmark—EXEBENCH—our model achieves 54.4% accuracy vs. 46.3% for LLM4Decompile and 37.5% for Nova; on REALTYPE, our model performs at least 95% better.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9197de72-87ca-4565-80ad-442a557a418eCited by top-tier papers1
Ask how each one uses itBuilds on23
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Debin: Predicting Debug Information in Stripped BinariesJingxuan He, Pesho Ivanov, Petar Tsankov, Veselin Raychev et al.CCS 2018 · 148 citations
- Neural reverse engineering of stripped binaries using augmented control flow graphsYaniv David, Uri Alon, Eran YahavOOPSLA 2020 · 83 citations
- OSPREY: Recovery of Variable and Data Structure via Probabilistic Analysis for Stripped BinaryZhuo Zhang, Yapeng Ye, Wei You, Guanhong Tao et al.S&P 2021 · 78 citations
Related papers
- QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly CodeHainan Fang, Yuanbo Wen, Jun Bi, Yihan Wang et al.NeurIPS 2025
- LLM4Decompile: Decompiling Binary Code with Large Language ModelsHanzhuo Tan, Qi Luo, Jing Li, Yuqun ZhangEMNLP 2024 · 31 citations
- HexT5: Unified Pre-Training for Stripped Binary Code Information InferenceJiaqi Xiong, Guoqiang Chen, Kejiang Chen, Han Gao et al.ASE 2023 · 8 citations
- Augmenting Decompiler Output with Learned Variable Names and TypesQibin Chen, Jeremy Lacomis, Edward J. Schwartz, Claire Le Goues et al.USENIX Security 2022
- SK2Decompile: LLM-based Two-Phase Binary Decompilation from Skeleton to SkinHanzhuo Tan, Weihao Li, Xiaolong Tian, Siyi Wang et al.ICLR 2026 · 10 citations
