Practical and Effective Code Watermarking for Large Language Models
Zhimeng Guo, Minhao Cheng
Abstract
The rapid advancement of Large Language Models (LLMs) in code generation has raised significant attribution and intellectual property concerns. Code watermarking offers a potential solution but faces unique challenges due to programming languages' strict syntactic constraints and semantic requirements. To address these challenges, we introduce ACW (AST-guided Code Watermarking), a novel adaptive framework that leverages Abstract Syntax Tree (AST) analysis during training to learn watermark embedding strategies. Our framework identifies substitutable code components and strategically biases token selections to embed watermarks. We also propose a novel sampling scheme that distributes tokens between green/red lists according to semantic context, ensuring statistical distinguishability while preserving code functionality. Extensive experiments demonstrate that ACW achieves a significant improvement in watermark detection accuracy compared to existing methods, with negligible impact on code functionality. This adaptive framework offers a promising solution for effective and practical code watermarking in the age of LLMs. Our code is available at: https://github.com/TimeLovercc/code-watermark.
Recent research has explored techniques like entropy-based methods and the utilization of variable type information to embed watermarks while maintaining type safety [19,9]. However, a significant limitation of these approaches lies in their detection phase, which often necessitates access to the
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3e2a9be-daa6-40a6-a9bf-7dad0681c57eBuilds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting et al.NeurIPS 2023 · 657 citations
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 312 citations
Related papers
- Adaptive Code Watermarking Through Reinforcement LearningZhimeng Guo, Huaisheng Zhu, Siyuan Xu, Hangfan Zhang et al.ICML 2026 · 73 citations
- CodeGenGuard: A Watermark for Code Generation ModelsBorui Yang, Mingxuan Ma, Liyao Xiang, Nan Chen et al.ICLR 2026
- From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language ModelsYidan Wang, Yubing Ren, Yanan Cao, Binxing FangACL 2025 · 4 citations
- AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language ModelsYue Li, Xin Yi, Dongsheng Shi, Yongyi Cui et al.KDD 2026 · 1 citation
- Who Wrote this Code? Watermarking for Code GenerationTaehyun Lee, Seokhee Hong, Jaewoo Ahn, Ilgee Hong et al.ACL 2024 · 36 citations
