FidelityGPT: Correcting Decompilation Distortions with Retrieval Augmented Generation
Zhiping Zhou, Xiaohong Li, Ruitao Feng, Yao Zhang, Yuekang Li, Wenbu Feng, Yunqian Wang, Yuqing Li
Abstract
Decompilation is a crucial technique that converts machine code into a human-readable format, facilitating analysis and debugging in the absence of source code. However, this process is hindered by fidelity issues, which can significantly impair the readability and accuracy of the decompiled output. Existing approaches partially addressed these, such as variable renaming and structural simplification, but typically fail to provide adequate detection and correction, especially in complex but practical closed-source binary scenarios. To address this, we introduce FidelityGPT, a novel framework to improve the accuracy and readability of decompiled code by systematically detecting and correcting discrepancies between decompiled code and its original source. FidelityGPT defines distortion prompt templates tailored to closed-source environments and incorporates Retrieval-Augmented Generation (RAG) with a dynamic semantic intensity algorithm. The algorithm identifies distorted lines based on semantic intensity, retrieving similar code from a database. Additionally, a variable dependency algorithm is designed to overcome the limitations of long-context inputs by analyzing redundant variables through their dependencies and integrating redundant variable names into prompt context. These combined techniques establish FidelityGPT as the first framework capable of effectively addressing decompilation distortion issues in LLM-based decompilation optimization. We evaluated FidelityGPT on 620 function pairs from a binary similarity benchmark, achieving an average detection accuracy of 89% and a precision of 83%. Compared to the current state-of-the-art model, DeGPT, which achieved an average Fix Rate (FR) of 83% and an average Corrected Fix Rate (CFR) of 37%, FidelityGPT demonstrated superior performance. With an average FR of 94% and an average CFR of 64%, FidelityGPT significantly improves both accuracy and readability, underscoring its effectiveness in enhancing decompilation and its potential to drive advancements in reverse engineering.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Oxidizer: Toward Concise and High-fidelity Rust DecompilationYibo Liu, Zion Leonahenahe Basque, Arvind S. Raj, Chavin Udomwongsa et al.S&P 2026 · 1 citation
- No More Translation at Runtime: LLM-Empowered Static Binary TranslationZhibo Liu, Huaijin Wang, Wai Kin Wong, Daoyuan Wu et al.EuroSys 2026 · 1 citation
Builds on18
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- The Power of Noise: Redefining Retrieval for RAG SystemsFlorin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice et al.SIGIR 2024 · 212 citations
- How far we have come: testing decompilation correctness of C decompilersZhibo Liu, Shuai WangISSTA 2020 · 53 citations
- LLM4Decompile: Decompiling Binary Code with Large Language ModelsHanzhuo Tan, Qi Luo, Jing Li, Yuqun ZhangEMNLP 2024 · 31 citations
- "Len or index or count, anything but v1": Predicting Variable Names in Decompilation Output with Transfer LearningKuntal Kumar Pal, Ati Priya Bajaj, Pratyay Banerjee, Audrey Dutcher et al.S&P 2024 · 30 citations
Related papers
- BinRAG: An RAG-Based Decompilation Framework Fusing Name Prediction and Calling ContextWai Kin Wong, Daoyuan Wu, Zhibo Liu, Huaijin Wang et al.ISSTA 2026
- DeGPT: Optimizing Decompiler Output with LLMPeiwei Hu, Ruigang Liang, Kai ChenNDSS 2024
- PseudoFix: Refactoring Distorted Structures in Decompiled C PseudocodeGangyang Li, Xiuwei Shang, Shaoyin Cheng, Junqi Zhang et al.ASE 2025 · 1 citation
- Between Lines of Code: Unraveling the Distinct Patterns of Machine and Human ProgrammersYuling Shi, Hongyu Zhang, Chengcheng Wan, Xiaodong GuICSE 2025 · 9 citations
- Learn-to-Distance: Distance Learning for Detecting LLM-Generated TextHongyi Zhou, Jin Zhu, Kai Ye, Ying Yang et al.ICLR 2026 · 10 citations
