SrcMarker: Dual-Channel Source Code Watermarking via Scalable Code Transformations
Borui Yang, Wei Li, Liyao Xiang, Bo Li
摘要
The expansion of the open source community and the rise of large language models have raised ethical and security concerns on the distribution of source code, such as misconduct on copyrighted code, distributions without proper licenses, or misuse of the code for malicious purposes. Hence it is important to track the ownership of source code, in which watermarking is a major technique. Yet, drastically different from natural languages, source code watermarking requires far stricter and more complicated rules to ensure the readability as well as the functionality of the source code. Hence we introduce SrcMarker, a watermarking system to unobtrusively encode ID bitstrings into source code, without affecting the usage and semantics of the code. To this end, SrcMarker performs transformations on an AST-based intermediate representation that enables unified transformations across different programming languages. The core of the system utilizes learning-based embedding and extraction modules to select rule-based transformations for watermarking. In addition, a novel feature-approximation technique is designed to tackle the inherent non-differentiability of rule selection, thus seamlessly integrating the rule-based transformations and learning-based networks into an interconnected system to enable end-to-end training. Extensive experiments demonstrate the superiority of SrcMarker over existing methods in various watermarking requirements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Waterfall: Scalable Framework for Robust Text Watermarking and Provenance for LLMsGregory Kang Ruey Lau, Xinyuan Niu, Hieu Dao, Jiangwei Chen 等EMNLP 2024 · 被引用 3 次
- CLMTracing: Black-box User-level Watermarking for Code Language Model TracingBoyu Zhang, Ping He, Tianyu Du, Xuhong Zhang 等EMNLP 2025 · 被引用 1 次
- Practical and Effective Code Watermarking for Large Language ModelsZhimeng Guo, Minhao ChengNeurIPS 2025 · 被引用 1 次
- CodeGenGuard: A Watermark for Code Generation ModelsBorui Yang, Mingxuan Ma, Liyao Xiang, Nan Chen 等ICLR 2026
- DeCoMa: Detecting and Purifying Code Dataset Watermarks through Dual Channel Code AbstractionYuan Xiao, Yuchen Chen, Shiqing Ma, Haocheng Huang 等ISSTA 2025
它引用的顶会 Paper15
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas 等USENIX Security 2018 · 被引用 832 次
- Entangled Watermarks as a Defense against Model ExtractionHengrui Jia, Christopher A. Choquette-Choo, Varun Chandrasekaran, Nicolas PapernotUSENIX Security 2021 · 被引用 287 次
- Adversarial Watermarking Transformer: Towards Tracing Text Provenance with Data HidingSahar Abdelnabi, Mario FritzS&P 2021 · 被引用 210 次
相关 Paper
- DuCodeMark: Dual-Purpose Code Dataset Watermarking via Style-Aware Watermark-Poison DesignYuchen Chen, Yuan Xiao, Chunrong Fang, Zhenyu Chen 等FSE 2026
- CodeMark: Imperceptible Watermarking for Code Datasets against Neural Code Completion ModelsZhensu Sun, Xiaoning Du, Fu Song, Li LiFSE 2023 · 被引用 34 次
- StealthInk: A Multi-bit and Stealthy Watermark for Large Language ModelsYa Jiang, Chuxiong Wu, Massieh Kordi Boroujeny, Brian L. Mark 等ICML 2025
- Provably Robust Multi-bit Watermarking for AI-generated TextWenjie Qu, Wengrui Zheng, Tianyang Tao, Dong Yin 等USENIX Security 2025
- Adaptive Code Watermarking Through Reinforcement LearningZhimeng Guo, Huaisheng Zhu, Siyuan Xu, Hangfan Zhang 等ICML 2026 · 被引用 73 次
