USENIX ATC2022顶会
Automatic Recovery of Fine-grained Compiler Artifacts at the Binary Level
Yufei Du, Ryan Court, Kevin Z. Snow, Fabian Monrose
摘要
Identifying a binary's compiler configuration enables developers and analysts to locate potential security issues caused by optimization side-effects, identify binary clones, and build compatible binary patches. Existing work focuses on identifying compiler family, version and optimization level of a binary using semantic features and deep learning techniques. Unfortunately, in practice, binaries are an amalgamation of objects and functions that can be compiled at different optimization levels with a variety of individual, fine-grained, optimizations that may be applied depending on the structure of the code. Hence, rather than recovering high-level artifacts, i.e., compiler family, version, and optimization level, we explore the recovery of individual, fine-grained, optimization passes for each function in a binary. To do so, we develop an approach using specially crafted features alongside intuitive and understandable machine learning models. Our evaluation on 15 popular open-source repositories shows that our approach compares favorable with the state-of-theart deep learning approach in compiler family, compiler version and optimization level identification. For finegrained optimization passes, our evaluation on 149,814 functions from 552 binaries in four popular open-source repositories shows that our approach achieves an average F-1 score of 92.1% for all optimization passes and an average F-1 score of 89.8% for optimization passes that could have negative impacts on security. Moreover, our approach includes experimental support for dynamic feature extraction via binary emulation, and our results shows that such features offer promising potential in improving the accuracy of optimization pass identification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Improving Security Tasks Using Compiler Provenance Information Recovered At the Binary-LevelYufei Du, Omar Alrawi, Kevin Z. Snow, Manos Antonakakis 等CCS 2023 · 被引用 8 次
- Enhancing Function Name Prediction using Votes-Based Name Tokenization and Multi-task LearningXiaoling Zhang, Zhengzi Xu, Shouguo Yang, Zhi Li 等FSE 2024 · 被引用 5 次
它引用的顶会 Paper3
- Automating Patching of Vulnerable Open-Source Software Versions in Application BinariesRuian Duan, Ashish Bijlani, Yang Ji, Omar Alrawi 等NDSS 2019 · 被引用 63 次
- Dead Store Elimination (Still) Considered HarmfulZhaomo Yang, Brian Johannesmeyer, Anders Trier Olesen, Sorin Lerner 等USENIX Security 2017 · 被引用 37 次
- Not so fast: understanding and mitigating negative impacts of compiler optimizations on code reuse gadget setsMichael D. Brown, Matthew Pruett, Robert Bigelow, Girish Mururu 等OOPSLA 2021 · 被引用 11 次
相关 Paper
- Improving Binary Code Similarity Transformer Models by Semantics-Driven Instruction DeemphasisXiangzhe Xu, Shiwei Feng, Yapeng Ye, Guangyu Shen 等ISSTA 2023 · 被引用 25 次
- DeepBinDiff: Learning Program-Wide Code Representations for Binary DiffingYue Duan, Xuezixiang Li, Jinghan Wang, Heng YinNDSS 2020
- When Coding Style Survives Compilation: De-anonymizing Programmers from Executable BinariesAylin Caliskan, Fabian Yamaguchi, Edwin Dauber, Richard E. Harang 等NDSS 2018 · 被引用 125 次
- A Deep Dive into Function Inlining and its Security Implications for ML-based Binary AnalysisOmar Abusabha, Jiyong Uhm, Tamer Abuhmed, Hyungjoon KooNDSS 2026 · 被引用 3 次
- Revisiting Optimization-Resilience Claims in Binary Diffing Tools: Insights from LLVM Peephole Optimization AnalysisXiaolei Ren, Mengfei Ren, Yu Lei, Jiang MingFSE 2025
