LLM4Decompile: Decompiling Binary Code with Large Language Models
Hanzhuo Tan, Qi Luo, Jing Li, Yuqun Zhang
摘要
Decompilation aims to convert binary code to high-level source code, but traditional tools like Ghidra often produce results that are difficult to read and execute. Motivated by the advancements in Large Language Models (LLMs), we propose LLM4Decompile, the first and largest open-source LLM series (1.3B to 33B) trained to decompile binary code. We optimize the LLM training process and introduce the LLM4Decompile-End models to decompile binary directly. The resulting models significantly outperform GPT-4o and Ghidra on the HumanEval and ExeBench benchmarks over 100% in terms of re-executability rate. Additionally, we improve the standard refinement approach to fine-tune the LLM4Decompile-Ref models, enabling them to effectively refine the decompiled code from Ghidra and achieve a further 16.2% improvement over the LLM4Decompile-End. LLM4Decompile 1 demonstrates the potential of LLMs to revolutionize binary code decompilation, delivering remarkable improvements in readability and executability while complementing conventional tools for optimal results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- ReSym: Harnessing LLMs to Recover Variable and Data Structure Symbols from Stripped BinariesDanning Xie, Zhuo Zhang, Nan Jiang, Xiangzhe Xu 等CCS 2024 · 被引用 21 次
- SK2Decompile: LLM-based Two-Phase Binary Decompilation from Skeleton to SkinHanzhuo Tan, Weihao Li, Xiaolong Tian, Siyi Wang 等ICLR 2026 · 被引用 10 次
- ShieldedCode: Learning Robust Representations for Virtual Machine Protected CodeMingqiao Mo, Yunlong Tan, Hao Zhang, Heng Zhang 等ICLR 2026 · 被引用 10 次
- WaDec: Decompiling WebAssembly Using Large Language ModelXinyu She, Yanjie Zhao, Haoyu WangASE 2024 · 被引用 9 次
- DecLLM: LLM-Augmented Recompilable Decompilation for Enabling Programmatic Use of Decompiled CodeWai Kin Wong, Daoyuan Wu, Huaijin Wang, Zongjie Li 等ISSTA 2025 · 被引用 8 次
它引用的顶会 Paper8
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- RetroWrite: Statically Instrumenting COTS Binaries for Fuzzing and SanitizationSushant Dinesh, Nathan Burow, Dongyan Xu, Mathias PayerS&P 2020 · 被引用 187 次
- DOBF: A Deobfuscation Pre-Training Objective for Programming LanguagesMarie-Anne Lachaux, Baptiste Rozière, Marc Szafraniec, Guillaume LampleNeurIPS 2021 · 被引用 174 次
相关 Paper
- DeGPT: Optimizing Decompiler Output with LLMPeiwei Hu, Ruigang Liang, Kai ChenNDSS 2024
- BinRAG: An RAG-Based Decompilation Framework Fusing Name Prediction and Calling ContextWai Kin Wong, Daoyuan Wu, Zhibo Liu, Huaijin Wang 等ISSTA 2026
- Can Large Language Models Understand Intermediate Representations in Compilers?Hailong Jiang, Jianfeng Zhu, Yao Wan, Bo Fang 等ICML 2025
- Enhancing LLMs with Staged Grouping and Dehallucination for Header File DecompositionYue Wang, Jiaxuan Sun, Yanzhen Zou, Bing XieASE 2025
- Idioms: A Simple and Effective Framework for Turbo-Charging Local Neural Decompilation with Well-Defined TypesLuke Dramko, Claire Le Goues, Edward J. SchwartzNDSS 2026 · 被引用 7 次
