Beyond Dataset Watermarking: Model-Level Copyright Protection for Code Summarization Models
Jiale Zhang, Haoxuan Li, Di Wu, Xiaobing Sun, Qinghua Lu, Guodong Long
摘要
Code Summarization Model (CSM) has been widely used in code production, such as online and web programming for PHP and Javascript. CSMs are essential tools in code production, enhancing software development efficiency and driving innovation in automated code analysis. However, CSMs face risks of exploitation by unauthorized users, particularly in an online environment where CSMs can be easily shared and disseminated. To address these risks, digital watermarks offer a promising solution by embedding imperceptible signatures within the models to assert copyright ownership and track unauthorized usage. Traditional watermarking for CSM copyright protection faces two main challenges: 1) dataset watermarking methods require separate design of triggers and watermark features based on the characteristics of different programming languages, which not only increases the computation complexity but also leads to a lack of generalization, 2) existing watermarks based on code style transformation are easily identifiable by automated detection, demonstrating poor concealment. To tackle these issues, we propose ModMark, a novel model-level digital watermark embedding method. Specifically, by fine-tuning the tokenizer, ModMark achieves cross-language generalization while reducing the complexity of watermark design. Moreover, we employ noise injection techniques to effectively prevent trigger detection. Experimental results show that our method can achieve 100% watermark verification rate across various programming languages' CSMs, and the concealment and effectiveness of ModMark can also be guaranteed. Our codes and datasets are available at https://github.com/Ocreatedin/ModMark . CCS Concepts • Security and privacy → Software and application security.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CodeGenGuard: A Watermark for Code Generation ModelsBorui Yang, Mingxuan Ma, Liyao Xiang, Nan Chen 等ICLR 2026
- PuzzleMark: Implicit Jigsaw Learning for Robust Code Dataset Watermarking in Neural Code Completion ModelsHaocheng Huang, Yuchen Chen, Weisong Sun, Peizhuo Lv 等FSE 2026
它引用的顶会 Paper11
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- CoProtector: Protect Open-Source Code against Unauthorized Training Usage with Data PoisoningZhensu Sun, Xiaoning Du, Fu Song, Mingze Ni 等WWW 2022 · 被引用 95 次
- Training-free Lexical Backdoor Attacks on Language ModelsYujin Huang, Terry Yue Zhuo, Qiongkai Xu, Han Hu 等WWW 2023 · 被引用 56 次
- Practitioners' Expectations on Automated Code Comment GenerationXing Hu, Xin Xia, David Lo, Zhiyuan Wan 等ICSE 2022 · 被引用 50 次
相关 Paper
- CodeMark: Imperceptible Watermarking for Code Datasets against Neural Code Completion ModelsZhensu Sun, Xiaoning Du, Fu Song, Li LiFSE 2023 · 被引用 34 次
- DuCodeMark: Dual-Purpose Code Dataset Watermarking via Style-Aware Watermark-Poison DesignYuchen Chen, Yuan Xiao, Chunrong Fang, Zhenyu Chen 等FSE 2026
- DeCoMa: Detecting and Purifying Code Dataset Watermarks through Dual Channel Code AbstractionYuan Xiao, Yuchen Chen, Shiqing Ma, Haocheng Huang 等ISSTA 2025
- SrcMarker: Dual-Channel Source Code Watermarking via Scalable Code TransformationsBorui Yang, Wei Li, Liyao Xiang, Bo LiS&P 2024 · 被引用 21 次
- CLMTracing: Black-box User-level Watermarking for Code Language Model TracingBoyu Zhang, Ping He, Tianyu Du, Xuhong Zhang 等EMNLP 2025 · 被引用 1 次
