Protecting Intellectual Property of Large Language Model-Based Code Generation APIs via Watermarks
Zongjie Li, Chaozheng Wang, Shuai Wang, Cuiyun Gao
Abstract
The rise of large language model-based code generation (LLCG) has enabled various commercial services and APIs. Training LLCG models is often expensive and time-consuming, and the training data are often large-scale and even inaccessible to the public. As a result, the risk of intellectual property (IP) theft over the LLCG models (e.g., via imitation attacks) has been a serious concern. In this paper, we propose the first watermark (WM) technique to protect LLCG APIs from remote imitation attacks. Our proposed technique is based on replacing tokens in an LLCG output with their "synonyms" available in the programming language. A WM is thus defined as the stealthily tweaked distribution among token synonyms in LLCG outputs. We design six WM schemes (instantiated into over 30 WM passes) which rely on conceptually distinct token synonyms available in programming languages. Moreover, to check the IP of a suspicious model (decide if it is stolen from our protected LLCG API), we propose a statistical tests-based procedure that can directly check a remote, suspicious LLCG API.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8a93e70b-404d-4ca0-94fb-be69b7092386Cited by top-tier papers19
- REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language ModelsRuisi Zhang, Shehzeen Samarah Hussain, Paarth Neekhara, Farinaz KoushanfarUSENIX Security 2024 · 88 citations
- Adaptive Code Watermarking Through Reinforcement LearningZhimeng Guo, Huaisheng Zhu, Siyuan Xu, Hangfan Zhang et al.ICML 2026 · 73 citations
- Watermarking Makes Language Models RadioactiveTom Sander, Pierre Fernandez, Alain Durmus, Matthijs Douze et al.NeurIPS 2024 · 68 citations
- Who Wrote this Code? Watermarking for Code GenerationTaehyun Lee, Seokhee Hong, Jaewoo Ahn, Ilgee Hong et al.ACL 2024 · 36 citations
- LLM Whisperer: An Inconspicuous Attack to Bias LLM ResponsesWeiran Lin, Anna Gerchanovsky, Omer Akgul, Lujo Bauer et al.CHI 2025 · 23 citations
Related papers
- Protecting Intellectual Property of Language Generation APIs with Lexical WatermarkXuanli He, Qiongkai Xu, Lingjuan Lyu, Fangzhao Wu et al.AAAI 2022 · 124 citations
- Protecting Language Generation Models via Invisible WatermarkingXuandong Zhao, Yu-Xiang Wang, Lei LiICML 2023 · 117 citations
- WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation WatermarksAnudeex Shetty, Qiongkai Xu, Jey Han LauACL 2025 · 7 citations
- CATER: Intellectual Property Protection on Text Generation APIs via Conditional WatermarksXuanli He, Qiongkai Xu, Yi Zeng, Lingjuan Lyu et al.NeurIPS 2022 · 106 citations
- RAG-WM: An Efficient Black-Box Watermarking Approach for Retrieval-Augmented Generation of Large Language ModelsPeizhuo Lv, Mengjie Sun, Hao Wang, XiaoFeng Wang et al.CCS 2025
