Detecting Automatic Software Plagiarism via Token Sequence Normalization
Timur Saglam, Moritz Brödel, Larissa Schmid, Sebastian Hahner
Abstract
While software plagiarism detectors have been used for decades, the assumption that evading detection requires programming proficiency is challenged by the emergence of automated plagiarism generators. These generators enable effortless obfuscation attacks, exploiting vulnerabilities in existing detectors by inserting statements to disrupt the matching of related programs. Thus, we present a novel, language-independent defense mechanism that leverages program dependence graphs, rendering such attacks infeasible. We evaluate our approach with multiple real-world datasets and show that it defeats plagiarism generators by offering resilience against automated obfuscation while maintaining a low rate of false positives.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2af9bf9-b2f5-405a-9d69-f8fbd14b73bcCited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Authorship attribution of source code: a language-agnostic approach and applicability in software engineeringEgor Bogomolov, Vladimir Kovalenko, Yurii Rebryk, Alberto Bacchelli et al.FSE 2021 · 45 citations
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting et al.NeurIPS 2023 · 657 citations
- AI Wrote My Paper and All I Got was This False Negative:* Measuring the Efficacy of Commercial AI Text DetectorsSeth Layton, Bernardo B. P. Medeiros, Kevin R. B. Butler, Patrick TraynorS&P 2026 · 3 citations
- VULGEN: Realistic Vulnerability Generation Via Pattern Mining and Deep LearningYu Nong, Yuzhe Ou, Michael Pradel, Feng Chen et al.ICSE 2023 · 32 citations
- Detecting Data Poisoning in Code Generation LLMs via Black-Box, Vulnerability-Oriented ScanningShenao Yan, Shan Jin, Shimaa Ahmed, Sunpreet Singh Arora et al.CCS 2026
