When Coding Style Survives Compilation: De-anonymizing Programmers from Executable Binaries
Aylin Caliskan, Fabian Yamaguchi, Edwin Dauber, Richard E. Harang, Konrad Rieck, Rachel Greenstadt, Arvind Narayanan
摘要
The ability to identify authors of computer programs based on their coding style is a direct threat to the privacy and anonymity of programmers. While recent work found that source code can be attributed to authors with high accuracy, attribution of executable binaries appears to be much more difficult. Many distinguishing features present in source code, e.g. variable names, are removed in the compilation process, and compiler optimization may alter the structure of a program, further obscuring features that are known to be useful in determining authorship. We examine programmer de-anonymization from the standpoint of machine learning, using a novel set of features that include ones obtained by decompiling the executable binary to source code. We adapt a powerful set of techniques from the domain of source code authorship attribution along with stylistic representations embedded in assembly, resulting in successful de-anonymization of a large set of programmers. We evaluate our approach on data from the Google Code Jam, obtaining attribution accuracy of up to 96% with 100 and 83% with 600 candidate programmers. We present an executable binary authorship attribution approach, for the first time, that is robust to basic obfuscations, a range of compiler optimization settings, and binaries that have been stripped of their symbol tables. We perform programmer de-anonymization using both obfuscated binaries, and real-world code found "in the wild" in single-author GitHub repositories and the recently leaked Nulled.IO hacker forum. We show that programmers who would like to remain anonymous need to take extreme countermeasures to protect their privacy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- MemGuard: Defending against Black-Box Membership Inference Attacks via Adversarial ExamplesJinyuan Jia, Ahmed Salem, Michael Backes, Yang Zhang 等CCS 2019 · 被引用 464 次
- Debin: Predicting Debug Information in Stripped BinariesJingxuan He, Pesho Ivanov, Petar Tsankov, Veselin Raychev 等CCS 2018 · 被引用 148 次
- Misleading Authorship Attribution of Source Code using Adversarial LearningErwin Quiring, Alwin Maier, Konrad RieckUSENIX Security 2019 · 被引用 123 次
- Large-Scale and Language-Oblivious Code Authorship IdentificationMohammed Abuhamad, Tamer AbuHmed, Aziz Mohaisen, DaeHun NyangCCS 2018 · 被引用 102 次
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
相关 Paper
- Automatic Recovery of Fine-grained Compiler Artifacts at the Binary LevelYufei Du, Ryan Court, Kevin Z. Snow, Fabian MonroseUSENIX ATC 2022
- Enhancing Robustness of Code Authorship Attribution through Expert Feature KnowledgeXiaowei Guo, Cai Fu, Juan Chen, Hongle Liu 等ISSTA 2024 · 被引用 2 次
- Adversarial Authorship Attribution for DeobfuscationWanyue Zhai, Jonathan Rusert, Zubair Shafiq, Padmini SrinivasanACL 2022 · 被引用 7 次
- RoPGen: Towards Robust Code Authorship Attribution via Automatic Coding Style TransformationZhen Li, Qian (Guenevere) Chen, Chen Chen, Yayi Zou 等ICSE 2022 · 被引用 39 次
- An In-Depth Analysis of Disassembly on Full-Scale x86/x64 BinariesDennis Andriesse, Xi Chen, Victor van der Veen, Asia Slowinska 等USENIX Security 2016 · 被引用 162 次
