Enhancing Robustness of Code Authorship Attribution through Expert Feature Knowledge
Xiaowei Guo, Cai Fu, Juan Chen, Hongle Liu, Lansheng Han, Wenjin Li
Abstract
Code authorship attribution has been an interesting research problem for decades. Recent studies have revealed that existing methods for code authorship attribution suffer from weak robustness. Under the influence of small perturbations added by the attacker, the accuracy of the method will be greatly reduced. As of now, there is no code authorship attribution method capable of effectively handling such attacks. In this paper, we attribute the weak robustness of code authorship attribution methods to dataset bias and argue that this bias can be mitigated through adjustments to the feature learning strategy. We first propose a robust code authorship attribution feature combination framework, which is composed of only simple shallow neural network structures, and introduces controllability for the framework in the feature extraction by incorporating expert knowledge. Experiments show that the framework has significantly improved robustness over mainstream code authorship attribution methods, with an average drop of 23.4% (from 37.8% to 14.3%) in the success rate of targeted attacks and 25.9% (from 46.7% to 20.8%) in the success rate of untargeted attacks. At the same time, it can also achieve results comparable to mainstream code authorship attribution methods in terms of accuracy.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 596db26e-0c29-4144-a6cc-7daec880c97fRelated papers
- RoPGen: Towards Robust Code Authorship Attribution via Automatic Coding Style TransformationZhen Li, Qian (Guenevere) Chen, Chen Chen, Yayi Zou et al.ICSE 2022 · 39 citations
- Misleading Authorship Attribution of Source Code using Adversarial LearningErwin Quiring, Alwin Maier, Konrad RieckUSENIX Security 2019 · 123 citations
- Adversarial Robustness for CodePavol Bielik, Martin T. VechevICML 2020 · 101 citations
- Robin: A Novel Method to Produce Robust Interpreters for Deep Learning-Based Code ClassifiersZhen Li, Ruqian Zhang, Deqing Zou, Ning Wang et al.ASE 2023 · 4 citations
- Large-Scale and Language-Oblivious Code Authorship IdentificationMohammed Abuhamad, Tamer AbuHmed, Aziz Mohaisen, DaeHun NyangCCS 2018 · 102 citations
