DuCodeMark: Dual-Purpose Code Dataset Watermarking via Style-Aware Watermark-Poison Design
Yuchen Chen, Yuan Xiao, Chunrong Fang, Zhenyu Chen, Baowen Xu
摘要
The proliferation of large language models for code (CodeLMs) and open-source contributions has heightened concerns over unauthorized use of source code datasets. While watermarking provides a viable protection mechanism by embedding ownership signals, existing methods rely on detectable trigger-target patterns and are limited to source-code tasks, overlooking other scenarios such as decompilation tasks. In this paper, we propose DuCodeMark, a stealthy and robust dual-purpose watermarking method for code datasets that generalizes across both source-code tasks and decompilation tasks. DuCodeMark parses each code sample into an abstract syntax tree (AST), applies language-specific style transformations to construct stealthy trigger-target pairs, and injects repressible poisoned features into a subset of return-typed samples to enhance robustness against watermark removal or evasion. These features remain inactive during normal training but are activated upon watermark removal, degrading model performance. For verification, DuCodeMark employs a black-box method based on the independent-samples 𝑡-test. We conduct a comprehensive evaluation of DuCodeMark across 72 settings spanning two code tasks, two programming languages, three CodeLMs, and six decoding temperatures. The results demonstrate that it consistently achieves strong verifiability (𝑝 < 0.05), high stealthiness (suspicion rate ≤ 0.36), robustness against both watermark and poisoning attacks (recall ≤ 0.57), and a substantial drop in model performance upon watermark removal (Pass@1 drops by 28.6%), underscoring its practicality and resilience.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu 等ICLR 2023 · 被引用 234 次
- Multi-task Learning based Pre-trained Language Model for Code CompletionFang Liu, Ge Li, Yunfei Zhao, Zhi JinASE 2020 · 被引用 162 次
- Natural Attack for Pre-trained Models of CodeZhou Yang, Jieke Shi, Junda He, David LoICSE 2022 · 被引用 150 次
- CoProtector: Protect Open-Source Code against Unauthorized Training Usage with Data PoisoningZhensu Sun, Xiaoning Du, Fu Song, Mingze Ni 等WWW 2022 · 被引用 95 次
- Code Search based on Context-aware Code TranslationWeisong Sun, Chunrong Fang, Yuchen Chen, Guanhong Tao 等ICSE 2022 · 被引用 53 次
相关 Paper
- DeCoMa: Detecting and Purifying Code Dataset Watermarks through Dual Channel Code AbstractionYuan Xiao, Yuchen Chen, Shiqing Ma, Haocheng Huang 等ISSTA 2025
- CodeMark: Imperceptible Watermarking for Code Datasets against Neural Code Completion ModelsZhensu Sun, Xiaoning Du, Fu Song, Li LiFSE 2023 · 被引用 34 次
- SrcMarker: Dual-Channel Source Code Watermarking via Scalable Code TransformationsBorui Yang, Wei Li, Liyao Xiang, Bo LiS&P 2024 · 被引用 21 次
- PuzzleMark: Implicit Jigsaw Learning for Robust Code Dataset Watermarking in Neural Code Completion ModelsHaocheng Huang, Yuchen Chen, Weisong Sun, Peizhuo Lv 等FSE 2026
- CLMTracing: Black-box User-level Watermarking for Code Language Model TracingBoyu Zhang, Ping He, Tianyu Du, Xuhong Zhang 等EMNLP 2025 · 被引用 1 次
