Large-Scale and Language-Oblivious Code Authorship Identification
Mohammed Abuhamad, Tamer AbuHmed, Aziz Mohaisen, DaeHun Nyang
摘要
Eicient extraction of code authorship attributes is key for successful identiication. However, the extraction of such attributes is very challenging, due to various programming language speciics, the limited number of available code samples per author, and the average code lines per ile, among others. To this end, this work proposes a Deep Learning-based Code Authorship Identiication System (DL-CAIS) for code authorship attribution that facilitates large-scale, language-oblivious, and obfuscation-resilient code authorship identiication. The deep learning architecture adopted in this work includes TF-IDF-based deep representation using multiple Recurrent Neural Network (RNN) layers and fully-connected layers dedicated to authorship attribution learning. The deep representation then feeds into a random forest classiier for scalability to de-anonymize the author. Comprehensive experiments are conducted to evaluate DL-CAIS over the entire Google Code Jam (GCJ) dataset across all years (from 2008 to 2016) and over real-world code samples from 1987 public repositories on GitHub. The results of our work show the high accuracy despite requiring a smaller number of iles per author. Namely, we achieve an accuracy of 96% when experimenting with 1,600 authors for GCJ, and 94.38% for the real-world dataset for 745 C programmers. Our system also allows us to identify 8,903 authors, the largest-scale dataset used by far, with an accuracy of 92.3%. Moreover, our technique is resilient to language-speciics, and thus it can identify authors of four programming languages (e.g., C, C++, Java, and Python), and authors writing in mixed languages (e.g., Java/C++, Python/C++). Finally, our system is resistant to sophisticated obfuscation (e.g., using C Tigress) with an accuracy of 93.42% for a set of 120 authors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Misleading Authorship Attribution of Source Code using Adversarial LearningErwin Quiring, Alwin Maier, Konrad RieckUSENIX Security 2019 · 被引用 123 次
- DFD: Adversarial Learning-based Approach to Defend Against Website FingerprintingAhmed Abusnaina, Rhongho Jang, Aminollah Khormali, DaeHun Nyang 等INFOCOM 2020 · 被引用 51 次
- RoPGen: Towards Robust Code Authorship Attribution via Automatic Coding Style TransformationZhen Li, Qian (Guenevere) Chen, Chen Chen, Yayi Zou 等ICSE 2022 · 被引用 39 次
- SrcMarker: Dual-Channel Source Code Watermarking via Scalable Code TransformationsBorui Yang, Wei Li, Liyao Xiang, Bo LiS&P 2024 · 被引用 21 次
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
它引用的顶会 Paper1
相关 Paper
- DeMinify: Neural Variable Name Recovery and Type InferenceYi Li, Aashish Yadavally, Jiaxing Zhang, Shaohua Wang 等FSE 2023 · 被引用 6 次
- Enhancing Robustness of Code Authorship Attribution through Expert Feature KnowledgeXiaowei Guo, Cai Fu, Juan Chen, Hongle Liu 等ISSTA 2024 · 被引用 2 次
- DeciX: Explain Deep Learning Based Code Generation ApplicationsSimin Chen, Zexin Li, Wei Yang, Cong LiuFSE 2024 · 被引用 1 次
- Automated Website Fingerprinting through Deep LearningVera Rimmer, Davy Preuveneers, Marc Juarez, Tom van Goethem 等NDSS 2018 · 被引用 399 次
- A Context-based Automated Approach for Method Name Consistency Checking and SuggestionYi Li, Shaohua Wang, Tien N. NguyenICSE 2021 · 被引用 36 次
