RoPGen: Towards Robust Code Authorship Attribution via Automatic Coding Style Transformation
Zhen Li, Qian (Guenevere) Chen, Chen Chen, Yayi Zou, Shouhuai Xu
Abstract
Source code authorship attribution is an important problem often encountered in applications such as software forensics, bug fixing, and software quality analysis. Recent studies show that current source code authorship attribution methods can be compromised by attackers exploiting adversarial examples and coding style manipulation. This calls for robust solutions to the problem of code authorship attribution. In this paper, we initiate the study on making Deep Learning (DL)-based code authorship attribution robust. We propose an innovative framework called Robust coding style Patterns Generation (RoPGen), which essentially learns authors' unique coding style patterns that are hard for attackers to manipulate or imitate. The key idea is to combine data augmentation and gradient augmentation at the adversarial training phase. This effectively increases the diversity of training examples, generates meaningful perturbations to gradients of deep neural networks, and learns diversified representations of coding styles. We evaluate the effectiveness of RoPGen using four datasets of programs written in C, C++, and Java. Experimental results show that RoPGen can significantly improve the robustness of DL-based code authorship attribution, by respectively reducing 22.8% and 41.0% of the success rate of targeted and untargeted attacks on average. CCS CONCEPTS • Security and privacy → Software security engineering.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ca57ebb-a714-4ce1-af62-10666c0e9b72Cited by top-tier papers10
- An Extensive Study on Adversarial Attack against Pre-trained Models of CodeXiaohu Du, Ming Wen, Zichao Wei, Shangwen Wang et al.FSE 2023 · 21 citations
- SrcMarker: Dual-Channel Source Code Watermarking via Scalable Code TransformationsBorui Yang, Wei Li, Liyao Xiang, Bo LiS&P 2024 · 21 citations
- Two Sides of the Same Coin: Exploiting the Impact of Identifiers in Neural Code ComprehensionShuzheng Gao, Cuiyun Gao, Chaozheng Wang, Jun Sun et al.ICSE 2023 · 17 citations
- Attribution-guided Adversarial Code Prompt Generation for Code Completion ModelsXueyang Li, Guozhu Meng, Shangqing Liu, Lu Xiang et al.ASE 2024 · 5 citations
- ClassEval-T: Evaluating Large Language Models in Class-Level Code TranslationPengyu Xue, Linhao Wu, Zhen Yang, Chengyi Wang et al.ISSTA 2025 · 5 citations
Builds on15
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- Adversarial Robustness Against the Union of Multiple Perturbation ModelsPratyush Maini, Eric Wong, J. Zico KolterICML 2020 · 171 citations
- Adversarial examples for models of codeNoam Yefet, Uri Alon, Eran YahavOOPSLA 2020 · 162 citations
- Generating Adversarial Examples for Holding Robustness of Source Code Processing ModelsHuangzhao Zhang, Zhuo Li, Ge Li, Lei Ma et al.AAAI 2020 · 148 citations
Related papers
- Enhancing Robustness of Code Authorship Attribution through Expert Feature KnowledgeXiaowei Guo, Cai Fu, Juan Chen, Hongle Liu et al.ISSTA 2024 · 2 citations
- Misleading Authorship Attribution of Source Code using Adversarial LearningErwin Quiring, Alwin Maier, Konrad RieckUSENIX Security 2019 · 123 citations
- AACEGEN: Attention Guided Adversarial Code Example Generation for Deep Code ModelsZhong Li, Chong Zhang, Minxue Pan, Tian Zhang et al.ASE 2024 · 4 citations
- A Causal Learning Framework for Enhancing Robustness of Source Code ModelsJunyao Ye, Zhen Li, Xi Tang, Deqing Zou et al.FSE 2025
- Adversarial Robustness for CodePavol Bielik, Martin T. VechevICML 2020 · 101 citations
