Between Lines of Code: Unraveling the Distinct Patterns of Machine and Human Programmers
Yuling Shi, Hongyu Zhang, Chengcheng Wan, Xiaodong Gu
摘要
Large language models have catalyzed an unprece-dented wave in code generation. While achieving significant advances, they blur the distinctions between machine- and human-authored source code, causing integrity and authenticity issues of software artifacts. Previous methods such as DetectGPthave proven effective in discerning machine-generated texts, but they do not identify and harness the unique patterns of machine-generated code. Thus, its applicability falters when applied to code. In this paper, we carefully study the specific patterns that characterize machine- and human-authored code. Through a rigorous analysis of code attributes such as lexical diversity, conciseness, and naturalness, we expose unique patterns inherent to each source. We particularly notice that the syntactic segmentation of code is a critical factor in identifying its provenance. Based on our findings, we propose DetectCodeGPT, a novel method for detecting machine-generated code, which improves DetectGPT by capturing the distinct stylized patterns of code. Diverging from conventional techniques that depend on external LLMs for perturbations, DetectCodeGPT perturbs the code corpus by strategically inserting spaces and newlines, ensuring both efficacy and efficiency. Experiment results show that our approach significantly outperforms state-of-the-art techniques in detecting machine-generated code. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical DebuggingYuling Shi, Songsong Wang, Chengcheng Wan, Min Wang 等ICSE 2026 · 被引用 4 次
- SWE-Debate: Competitive Multi-Agent Debate for Software Issue ResolutionHan Li, Yuling Shi, Shaoxin Lin, Xiaodong Gu 等ICSE 2026 · 被引用 2 次
- Seeing Is Coding: On the Effectiveness of Vision Language Models in Code UnderstandingYuling Shi, Chaoxiang Xie, Zhensu Sun, Yeheng Chen 等ISSTA 2026 · 被引用 1 次
- Rethinking Code Complexity Through the Lens of Large Language ModelsChen Xie, Xiaodong Gu, Yuling Shi, Beijun ShenICML 2026
- CodeRipple: Wavelet-Based Detection of LLM-Generated CodeXingyu Yao, Zhendong Mao, Quan WangACL 2026
它引用的顶会 Paper14
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning 等ICML 2023 · 被引用 988 次
相关 Paper
- DualCodeDetect: Zero-Shot LLM-Generated Code Detection via Dual-Channel PerturbationZhengdao Li, Xiuwei Shang, Zhenkan Fu, Shikai Guo 等FSE 2026
- An Empirical Study on Automatically Detecting AI-Generated Source Code: How Far are We?Hyunjae Suh, Mahan Tafreshipour, Jiawei Li, Adithya Bhattiprolu 等ICSE 2025 · 被引用 2 次
- Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code RewritingTong Ye, Yangkai Du, Tengfei Ma, Lingfei Wu 等AAAI 2025 · 被引用 21 次
- Has My Code Been Stolen for Model Training? A Naturalness Based Approach to Code Contamination DetectionHaris Ali Khan, Yanjie Jiang, Qasim Umer, Yuxia Zhang 等FSE 2025 · 被引用 1 次
- Who Wrote this Code? Watermarking for Code GenerationTaehyun Lee, Seokhee Hong, Jaewoo Ahn, Ilgee Hong 等ACL 2024 · 被引用 36 次
