Disa: Accurate Learning-based Static Disassembly with Attentions
Peicheng Wang, Monika Santra, Mingyu Liu, Cong Sun, Dongrui Zeng, Gang Tan
摘要
For reverse engineering related security domains, such as vulnerability detection, malware analysis, and binary hardening, disassembly is crucial yet challenging. The fundamental challenge of disassembly is to identify instruction and function boundaries. Classic approaches rely on file-format assumptions and architecture-specific heuristics to guess the boundaries, resulting in incomplete and incorrect disassembly, especially when the binary is obfuscated. Recent advancements of disassembly have demonstrated that deep learning can improve both the accuracy and efficiency of disassembly. In this paper, we propose Disa, a new learning-based disassembly approach that uses the information of superset instructions over the multi-head self-attention to learn the instructions' correlations, thus being able to infer function entry-points and instruction boundaries. Disa can further identify instructions relevant to memory block boundaries to facilitate an advanced block-memory model based value-set analysis for an accurate control flow graph (CFG) generation. Our experiments show that Disa outperforms prior deep-learning disassembly approaches in function entry-point identification, especially achieving 9.1% and 13.2% F1-score improvement on binaries respectively obfuscated by the disassembly desynchronization technique and popular source-level obfuscator. By achieving an 18.5% improvement in the memory block precision, Disa generates more accurate CFGs with a 4.4% reduction in Average Indirect Call Targets (AICT) compared with the state-of-the-art heuristic-based approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- SOK: (State of) The Art of War: Offensive Techniques in Binary AnalysisYan Shoshitaishvili, Ruoyu Wang, Christopher Salls, Nick Stephens 等S&P 2016 · 被引用 1,085 次
- An In-Depth Analysis of Disassembly on Full-Scale x86/x64 BinariesDennis Andriesse, Xi Chen, Victor van der Veen, Asia Slowinska 等USENIX Security 2016 · 被引用 162 次
- Debin: Predicting Debug Information in Stripped BinariesJingxuan He, Pesho Ivanov, Petar Tsankov, Veselin Raychev 等CCS 2018 · 被引用 148 次
- PalmTree: Learning an Assembly Language Model for Instruction EmbeddingXuezixiang Li, Yu Qu, Heng YinCCS 2021 · 被引用 139 次
- Superset Disassembly: Statically Rewriting x86 Binaries Without HeuristicsErick Bauman, Zhiqiang Lin, Kevin W. HamlenNDSS 2018 · 被引用 112 次
相关 Paper
- DeepDi: Learning a Relational Graph Convolutional Network Model on Instructions for Fast and Accurate DisassemblySheng Yu, Yu Qu, Xunchao Hu, Heng YinUSENIX Security 2022
- XDA: Accurate, Robust Disassembly with Transfer LearningKexin Pei, Jonas Guan, David Williams-King, Junfeng Yang 等NDSS 2021
- DEEPVSA: Facilitating Value-set Analysis with Deep Learning for Postmortem Program AnalysisWenbo Guo, Dongliang Mu, Xinyu Xing, Min Du 等USENIX Security 2019 · 被引用 67 次
- NeuDep: neural binary memory dependence analysisKexin Pei, Dongdong She, Michael Wang, Scott Geng 等FSE 2022 · 被引用 8 次
- Analyzing Bytes: Pre-Disassembly Static Binary AnalysisHuan Nguyen, Soumyakant Priyadarshan, Chencheng Jiang, R. SekarPLDI 2026
