SK2Decompile: LLM-based Two-Phase Binary Decompilation from Skeleton to Skin
Hanzhuo Tan, Weihao Li, Xiaolong Tian, Siyi Wang, Jiaming Liu, Jing Li, Yuqun Zhang
摘要
Large Language Models (LLMs) have emerged as a promising approach for binary decompilation. However, the existing LLM-based decompilers are still somewhat limited in effectively presenting a program's source-level structure with its original identifiers. To mitigate this, we introduce SK 2 Decompile, a novel two-phase approach to decompile from the skeleton (semantic structure) to the skin (identifier) of programs. Specifically, we first apply a Structure Recovery model to translate a program's binary code to an Intermediate Representation (IR) as deriving the program's "skeleton", i.e., preserving control flow and data structures while obfuscating all identifiers with generic placeholders. We also apply reinforcement learning to reward the model for producing program structures that adhere to the syntactic and semantic rules expected by compilers. Second, we apply an Identifier Naming model to produce meaningful identifiers which reflect actual program semantics as deriving the program's "skin". We train the Identifier Naming model with a separate reinforcement learning objective that rewards the semantic similarity between its predictions and the reference code. Such a two-phase decompilation process facilitates advancing the correctness and readability of decompilation independently. Our evaluations indicate that SK 2 Decompile significantly outperforms the SOTA baselines, achieving 21.6% average re-executability rate gain over GPT-5-mini on the HumanEval dataset and 29.4% average R2I improvement over Idioms on the GitHub2025 benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based AgentsJiahong Xiang, Wenxiao He, Xihua Wang, Hongliang Tian 等ICSE 2026
- Empowering Autonomous Debugging Agents with Efficient Dynamic AnalysisJiahong Xiang, Xiaoyang Xu, Xiaopan Chu, Hongliang Tian 等FSE 2026
它引用的顶会 Paper20
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- DOBF: A Deobfuscation Pre-Training Objective for Programming LanguagesMarie-Anne Lachaux, Baptiste Rozière, Marc Szafraniec, Guillaume LampleNeurIPS 2021 · 被引用 174 次
- An extensive study on pre-trained models for program understanding and generationZhengran Zeng, Hanzhuo Tan, Haotian Zhang, Jing Li 等ISSTA 2022 · 被引用 142 次
- Neural reverse engineering of stripped binaries using augmented control flow graphsYaniv David, Uri Alon, Eran YahavOOPSLA 2020 · 被引用 83 次
相关 Paper
- LLM4Decompile: Decompiling Binary Code with Large Language ModelsHanzhuo Tan, Qi Luo, Jing Li, Yuqun ZhangEMNLP 2024 · 被引用 31 次
- BinRAG: An RAG-Based Decompilation Framework Fusing Name Prediction and Calling ContextWai Kin Wong, Daoyuan Wu, Zhibo Liu, Huaijin Wang 等ISSTA 2026
- DecLLM: LLM-Augmented Recompilable Decompilation for Enabling Programmatic Use of Decompiled CodeWai Kin Wong, Daoyuan Wu, Huaijin Wang, Zongjie Li 等ISSTA 2025 · 被引用 8 次
- PseudoFix: Refactoring Distorted Structures in Decompiled C PseudocodeGangyang Li, Xiuwei Shang, Shaoyin Cheng, Junqi Zhang 等ASE 2025 · 被引用 1 次
- Idioms: A Simple and Effective Framework for Turbo-Charging Local Neural Decompilation with Well-Defined TypesLuke Dramko, Claire Le Goues, Edward J. SchwartzNDSS 2026 · 被引用 7 次
