Source Code Foundation Models are Transferable Binary Analysis Knowledge Bases
Zian Su, Xiangzhe Xu, Ziyang Huang, Kaiyuan Zhang, Xiangyu Zhang
Abstract
Human-Oriented Binary Reverse Engineering (HOBRE) lies at the intersection of binary and source code, aiming to lift binary code to human-readable content relevant to source code, thereby bridging the binary-source semantic gap. Recent advancements in uni-modal code model pre-training, particularly in generative Source Code Foundation Models (SCFMs) and binary understanding models, have laid the groundwork for transfer learning applicable to HOBRE. However, existing approaches for HOBRE rely heavily on uni-modal models like SCFMs for supervised fine-tuning or general LLMs for prompting, resulting in sub-optimal performance. Inspired by recent progress in large multi-modal models, we propose that it is possible to harness the strengths of uni-modal code models from both sides to bridge the semantic gap effectively. In this paper, we introduce a novel probe-and-recover framework that incorporates a binary-source encoder-decoder model and black-box LLMs for binary analysis. Our approach leverages the pre-trained knowledge within SCFMs to synthesize relevant, symbol-rich code fragments as context. This additional context enables black-box LLMs to enhance recovery accuracy. We demonstrate significant improvements in zero-shot binary summarization and binary function name recovery, with a 10.3% relative gain in CHRF and a 16.7% relative gain in a GPT4-based metric for summarization, as well as a 6.7% and 7.4% absolute increase in token-level precision and recall for name recovery, respectively. These results highlight the effectiveness of our approach in automating and improving binary code analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- ReSym: Harnessing LLMs to Recover Variable and Data Structure Symbols from Stripped BinariesDanning Xie, Zhuo Zhang, Nan Jiang, Xiangzhe Xu et al.CCS 2024 · 21 citations
- SK2Decompile: LLM-based Two-Phase Binary Decompilation from Skeleton to SkinHanzhuo Tan, Weihao Li, Xiaolong Tian, Siyi Wang et al.ICLR 2026 · 10 citations
- Unleashing the Power of Generative Model in Recovering Variable Names from Stripped BinaryXiangzhe Xu, Zhuo Zhang, Zian Su, Ziyang Huang et al.NDSS 2025
- Retrofit: Continual Learning with Controlled Forgetting for Binary Security Detection and AnalysisYiling He, Junchi Lei, Hongyu She, Shuo Shao et al.USENIX Security 2026
- SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference AttacksKaiyuan Zhang, Siyuan Cheng, Hanxi Guo, Yuetian Chen et al.USENIX Security 2025
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- Multi-modal Learning for WebAssembly Reverse EngineeringHanxian Huang, Jishen ZhaoISSTA 2024 · 1 citation
- BinRAG: An RAG-Based Decompilation Framework Fusing Name Prediction and Calling ContextWai Kin Wong, Daoyuan Wu, Zhibo Liu, Huaijin Wang et al.ISSTA 2026
- HexT5: Unified Pre-Training for Stripped Binary Code Information InferenceJiaqi Xiong, Guoqiang Chen, Kejiang Chen, Han Gao et al.ASE 2023 · 8 citations
- MiSum: Multi-modality Heterogeneous Code Graph Learning for Multi-intent Binary Code SummarizationKangchen Zhu, Zhiliang Tian, Shangwen Wang, Weiguo Chen et al.FSE 2025 · 2 citations
- Hieronym: Leveraging Hierarchical Multi-Source Information for Function Renaming in Stripped BinaryXiaoling Zhang, Jian Sun, Dawei Wang, Chongyu Wang et al.CCS 2026
