SymLM: Predicting Function Names in Stripped Binaries via Context-Sensitive Execution-Aware Code Embeddings
Xin Jin, Kexin Pei, Jun Yeon Won, Zhiqiang Lin
摘要
Predicting function names in stripped binaries is an extremely useful but challenging task, as it requires summarizing the execution behavior and semantics of the function in human languages. Recently, there has been significant progress in this direction with machine learning. However, existing approaches fail to model the exhaustive function behavior and thus suffer from the poor generalizability to unseen binaries. To advance the state of the art, we present a function Symbol name prediction and binary Language Modeling (SymLM) framework, with a novel neural architecture that learns the comprehensive function semantics by jointly modeling the execution behavior of the calling context and instructions via a novel fusing encoder. We have evaluated SymLM with 1,431,169 binary functions from 27 popular open source projects, compiled with 4 optimizations (O0-O3) for 4 different architectures (i.e., x64, x86, ARM, and MIPS) and 4 obfuscations. SymLM outperforms the state-of-the-art function name prediction tools by up to 15.4%, 59.6%, and 35.0% in precision, recall, and F1 score, with significantly better generalizability and obfuscation resistance. Ablation studies also show that our design choices (e.g., fusing components of the calling context and execution behavior) substantially boost the performance of function name prediction. Finally, our case studies further demonstrate the practical use cases of SymLM in analyzing firmware images.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- FITS: Inferring Intermediate Taint Sources for Effective Vulnerability Analysis of IoT Device FirmwarePuzhuo Liu, Yaowen Zheng, Chengnian Sun, Chuan Qin 等ASPLOS 2023 · 被引用 21 次
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
- Source Code Foundation Models are Transferable Binary Analysis Knowledge BasesZian Su, Xiangzhe Xu, Ziyang Huang, Kaiyuan Zhang 等NeurIPS 2024 · 被引用 17 次
- SimLLM: Calculating Semantic Similarity in Code Summaries using a Large Language Model-Based ApproachXin Jin, Zhiqiang LinFSE 2024 · 被引用 8 次
- DecLLM: LLM-Augmented Recompilable Decompilation for Enabling Programmatic Use of Decompiled CodeWai Kin Wong, Daoyuan Wu, Huaijin Wang, Zongjie Li 等ISSTA 2025 · 被引用 8 次
它引用的顶会 Paper24
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- Asm2Vec: Boosting Static Representation Robustness for Binary Clone Search against Code Obfuscation and Compiler OptimizationSteven H. H. Ding, Benjamin C. M. Fung, Philippe CharlandS&P 2019 · 被引用 447 次
- Retrieval-based neural source code summarizationJian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun 等ICSE 2020 · 被引用 242 次
- Debin: Predicting Debug Information in Stripped BinariesJingxuan He, Pesho Ivanov, Petar Tsankov, Veselin Raychev 等CCS 2018 · 被引用 148 次
- Superset Disassembly: Statically Rewriting x86 Binaries Without HeuristicsErick Bauman, Zhiqiang Lin, Kevin W. HamlenNDSS 2018 · 被引用 112 次
相关 Paper
- Beyond Classification: Inferring Function Names in Stripped Binaries via Domain Adapted LLMsLinxi Jiang, Xin Jin, Zhiqiang LinNDSS 2025
- Hieronym: Leveraging Hierarchical Multi-Source Information for Function Renaming in Stripped BinaryXiaoling Zhang, Jian Sun, Dawei Wang, Chongyu Wang 等CCS 2026
- Enhancing Function Name Prediction using Votes-Based Name Tokenization and Multi-task LearningXiaoling Zhang, Zhengzi Xu, Shouguo Yang, Zhi Li 等FSE 2024 · 被引用 5 次
- XFL: Naming Functions in Binaries with Extreme Multi-label LearningJames Patrick-Evans, Moritz Dannehl, Johannes KinderS&P 2023
- A lightweight framework for function name reassignment based on large-scale stripped binariesHan Gao, Shaoyin Cheng, Yinxing Xue, Weiming ZhangISSTA 2021 · 被引用 46 次
