Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting
Tong Ye, Yangkai Du, Tengfei Ma, Lingfei Wu, Xuhong Zhang, Shouling Ji, Wenhai Wang
Abstract
Large Language Models (LLMs) have demonstrated remarkable proficiency in generating code. However, the misuse of LLM-generated (synthetic) code has raised concerns in both educational and industrial contexts, underscoring the urgent need for synthetic code detectors. Existing methods for detecting synthetic content are primarily designed for general text and struggle with code due to the unique grammatical structure of programming languages and the presence of numerous ``low-entropy'' tokens. Building on this, our work proposes a novel zero-shot synthetic code detector based on the similarity between the original code and its LLM-rewritten variants. Our method is based on the observation that differences between LLM-rewritten and original code tend to be smaller when the original code is synthetic. We utilize self-supervised contrastive learning to train a code similarity model and evaluate our approach on two synthetic code detection benchmarks. Our results demonstrate a significant improvement over existing SOTA synthetic content detectors, delivering notable gains in both performance and robustness on the APPS and MBPP benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 13b7b1fd-9eae-4a63-92c4-9949e0009719Cited by top-tier papers4
- Exploring ChatGPT's Capabilities on Vulnerability ManagementPeiyu Liu, Junming Liu, Lirong Fu, Kangjie Lu et al.USENIX Security 2024 · 52 citations
- Dynamic Bundling with Large Language Models for Zero-Shot Inference on Text-Attributed GraphsYusheng Zhao, Qixin Zhang, Xiao Luo, Weizhi Zhang et al.NeurIPS 2025 · 4 citations
- Chiseling Out Efficiency: Structured Skeleton Supervision for Efficient Code GenerationYu Yu, Zhihong Sun, Jia Li, Yao Wan et al.FSE 2026
- CodeChemist: Test-Time Scaling for Low-Resource Code Generation via Functional Knowledge TransferKaiXin Wang, Tianlin Li, Xiaoyu Zhang, Aishan Liu et al.ICML 2026
Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
Related papers
- DualCodeDetect: Zero-Shot LLM-Generated Code Detection via Dual-Channel PerturbationZhengdao Li, Xiuwei Shang, Zhenkan Fu, Shikai Guo et al.FSE 2026
- DeTeCtive: Detecting AI-generated Text via Multi-Level Contrastive LearningXun Guo, Yongxin He, Shan Zhang, Ting Zhang et al.NeurIPS 2024 · 100 citations
- Detecting Semantic Clones of Unseen FunctionalityKonstantinos Kitsios, Francesco Sovrano, Earl T. Barr, Alberto BacchelliASE 2025 · 1 citation
- Zero-Shot Detection of LLM-Generated Text via Implicit Reward ModelRunheng Liu, Heyan Huang, Xingchen Xiao, Zhijing WuNeurIPS 2025 · 7 citations
- An Empirical Study to Evaluate AIGC Detectors on Code ContentJian Wang, Shangqing Liu, Xiaofei Xie, Yi LiASE 2024 · 4 citations
