Beyond Dataset Watermarking: Model-Level Copyright Protection for Code Summarization Models
Jiale Zhang, Haoxuan Li, Di Wu, Xiaobing Sun, Qinghua Lu, Guodong Long
Abstract
Code Summarization Model (CSM) has been widely used in code production, such as online and web programming for PHP and Javascript. CSMs are essential tools in code production, enhancing software development efficiency and driving innovation in automated code analysis. However, CSMs face risks of exploitation by unauthorized users, particularly in an online environment where CSMs can be easily shared and disseminated. To address these risks, digital watermarks offer a promising solution by embedding imperceptible signatures within the models to assert copyright ownership and track unauthorized usage. Traditional watermarking for CSM copyright protection faces two main challenges: 1) dataset watermarking methods require separate design of triggers and watermark features based on the characteristics of different programming languages, which not only increases the computation complexity but also leads to a lack of generalization, 2) existing watermarks based on code style transformation are easily identifiable by automated detection, demonstrating poor concealment. To tackle these issues, we propose ModMark, a novel model-level digital watermark embedding method. Specifically, by fine-tuning the tokenizer, ModMark achieves cross-language generalization while reducing the complexity of watermark design. Moreover, we employ noise injection techniques to effectively prevent trigger detection. Experimental results show that our method can achieve 100% watermark verification rate across various programming languages' CSMs, and the concealment and effectiveness of ModMark can also be guaranteed. Our codes and datasets are available at https://github.com/Ocreatedin/ModMark . CCS Concepts • Security and privacy → Software and application security.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2f1056c-4bf3-4297-b551-567c2b9d0dd6Cited by top-tier papers2
- CodeGenGuard: A Watermark for Code Generation ModelsBorui Yang, Mingxuan Ma, Liyao Xiang, Nan Chen et al.ICLR 2026
- PuzzleMark: Implicit Jigsaw Learning for Robust Code Dataset Watermarking in Neural Code Completion ModelsHaocheng Huang, Yuchen Chen, Weisong Sun, Peizhuo Lv et al.FSE 2026
Builds on11
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- CoProtector: Protect Open-Source Code against Unauthorized Training Usage with Data PoisoningZhensu Sun, Xiaoning Du, Fu Song, Mingze Ni et al.WWW 2022 · 95 citations
- Training-free Lexical Backdoor Attacks on Language ModelsYujin Huang, Terry Yue Zhuo, Qiongkai Xu, Han Hu et al.WWW 2023 · 56 citations
- Practitioners' Expectations on Automated Code Comment GenerationXing Hu, Xin Xia, David Lo, Zhiyuan Wan et al.ICSE 2022 · 50 citations
Related papers
- CodeMark: Imperceptible Watermarking for Code Datasets against Neural Code Completion ModelsZhensu Sun, Xiaoning Du, Fu Song, Li LiFSE 2023 · 34 citations
- DuCodeMark: Dual-Purpose Code Dataset Watermarking via Style-Aware Watermark-Poison DesignYuchen Chen, Yuan Xiao, Chunrong Fang, Zhenyu Chen et al.FSE 2026
- DeCoMa: Detecting and Purifying Code Dataset Watermarks through Dual Channel Code AbstractionYuan Xiao, Yuchen Chen, Shiqing Ma, Haocheng Huang et al.ISSTA 2025
- SrcMarker: Dual-Channel Source Code Watermarking via Scalable Code TransformationsBorui Yang, Wei Li, Liyao Xiang, Bo LiS&P 2024 · 21 citations
- CLMTracing: Black-box User-level Watermarking for Code Language Model TracingBoyu Zhang, Ping He, Tianyu Du, Xuhong Zhang et al.EMNLP 2025 · 1 citation
