SrcMarker: Dual-Channel Source Code Watermarking via Scalable Code Transformations
Borui Yang, Wei Li, Liyao Xiang, Bo Li
Abstract
The expansion of the open source community and the rise of large language models have raised ethical and security concerns on the distribution of source code, such as misconduct on copyrighted code, distributions without proper licenses, or misuse of the code for malicious purposes. Hence it is important to track the ownership of source code, in which watermarking is a major technique. Yet, drastically different from natural languages, source code watermarking requires far stricter and more complicated rules to ensure the readability as well as the functionality of the source code. Hence we introduce SrcMarker, a watermarking system to unobtrusively encode ID bitstrings into source code, without affecting the usage and semantics of the code. To this end, SrcMarker performs transformations on an AST-based intermediate representation that enables unified transformations across different programming languages. The core of the system utilizes learning-based embedding and extraction modules to select rule-based transformations for watermarking. In addition, a novel feature-approximation technique is designed to tackle the inherent non-differentiability of rule selection, thus seamlessly integrating the rule-based transformations and learning-based networks into an interconnected system to enable end-to-end training. Extensive experiments demonstrate the superiority of SrcMarker over existing methods in various watermarking requirements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18671911-9978-41df-8856-3e7765e3fa85Cited by top-tier papers7
- Waterfall: Scalable Framework for Robust Text Watermarking and Provenance for LLMsGregory Kang Ruey Lau, Xinyuan Niu, Hieu Dao, Jiangwei Chen et al.EMNLP 2024 · 3 citations
- CLMTracing: Black-box User-level Watermarking for Code Language Model TracingBoyu Zhang, Ping He, Tianyu Du, Xuhong Zhang et al.EMNLP 2025 · 1 citation
- Practical and Effective Code Watermarking for Large Language ModelsZhimeng Guo, Minhao ChengNeurIPS 2025 · 1 citation
- CodeGenGuard: A Watermark for Code Generation ModelsBorui Yang, Mingxuan Ma, Liyao Xiang, Nan Chen et al.ICLR 2026
- DeCoMa: Detecting and Purifying Code Dataset Watermarks through Dual Channel Code AbstractionYuan Xiao, Yuchen Chen, Shiqing Ma, Haocheng Huang et al.ISSTA 2025
Builds on15
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas et al.USENIX Security 2018 · 832 citations
- Entangled Watermarks as a Defense against Model ExtractionHengrui Jia, Christopher A. Choquette-Choo, Varun Chandrasekaran, Nicolas PapernotUSENIX Security 2021 · 287 citations
- Adversarial Watermarking Transformer: Towards Tracing Text Provenance with Data HidingSahar Abdelnabi, Mario FritzS&P 2021 · 210 citations
Related papers
- DuCodeMark: Dual-Purpose Code Dataset Watermarking via Style-Aware Watermark-Poison DesignYuchen Chen, Yuan Xiao, Chunrong Fang, Zhenyu Chen et al.FSE 2026
- CodeMark: Imperceptible Watermarking for Code Datasets against Neural Code Completion ModelsZhensu Sun, Xiaoning Du, Fu Song, Li LiFSE 2023 · 34 citations
- StealthInk: A Multi-bit and Stealthy Watermark for Large Language ModelsYa Jiang, Chuxiong Wu, Massieh Kordi Boroujeny, Brian L. Mark et al.ICML 2025
- Provably Robust Multi-bit Watermarking for AI-generated TextWenjie Qu, Wengrui Zheng, Tianyang Tao, Dong Yin et al.USENIX Security 2025
- Adaptive Code Watermarking Through Reinforcement LearningZhimeng Guo, Huaisheng Zhu, Siyuan Xu, Hangfan Zhang et al.ICML 2026 · 73 citations
