Authorship attribution of source code: a language-agnostic approach and applicability in software engineering
Egor Bogomolov, Vladimir Kovalenko, Yurii Rebryk, Alberto Bacchelli, Timofey Bryksin
摘要
Authorship attribution (i.e., determining who is the author of a piece of source code) is an established research topic. State-of-the-art results for the authorship attribution problem look promising for the software engineering field, where they could be applied to detect plagiarized code and prevent legal issues. With this article, we first introduce a new language-agnostic approach to authorship attribution of source code. Then, we discuss limitations of existing synthetic datasets for authorship attribution, and propose a data collection approach that delivers datasets that better reflect aspects important for potential practical use in software engineering. Finally, we demonstrate that high accuracy of authorship attribution models on existing datasets drastically drops when they are evaluated on more realistic data. We outline next steps for the design and evaluation of authorship attribution models that could bring the research efforts closer to practical use for software engineering.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Malla: Demystifying Real-world Large Language Model Integrated Malicious ServicesZilong Lin, Jian Cui, Xiaojing Liao, XiaoFeng WangUSENIX Security 2024 · 被引用 49 次
- RoPGen: Towards Robust Code Authorship Attribution via Automatic Coding Style TransformationZhen Li, Qian (Guenevere) Chen, Chen Chen, Yayi Zou 等ICSE 2022 · 被引用 39 次
- Robin: A Novel Method to Produce Robust Interpreters for Deep Learning-Based Code ClassifiersZhen Li, Ruqian Zhang, Deqing Zou, Ning Wang 等ASE 2023 · 被引用 4 次
相关 Paper
- Misleading Authorship Attribution of Source Code using Adversarial LearningErwin Quiring, Alwin Maier, Konrad RieckUSENIX Security 2019 · 被引用 123 次
- Detecting Automatic Software Plagiarism via Token Sequence NormalizationTimur Saglam, Moritz Brödel, Larissa Schmid, Sebastian HahnerICSE 2024 · 被引用 6 次
- Enhancing Robustness of Code Authorship Attribution through Expert Feature KnowledgeXiaowei Guo, Cai Fu, Juan Chen, Hongle Liu 等ISSTA 2024 · 被引用 2 次
- A Girl Has A Name: Detecting Authorship ObfuscationAsad Mahmood, Zubair Shafiq, Padmini SrinivasanACL 2020 · 被引用 1 次
- Adversarial Authorship Attribution for DeobfuscationWanyue Zhai, Jonathan Rusert, Zubair Shafiq, Padmini SrinivasanACL 2022 · 被引用 7 次
