Reliable Model Watermarking: Defending against Theft without Compromising on Evasion
Hongyu Zhu, Sichu Liang, Wentao Hu, Fangqi Li, Ju Jia, Shi-Lin Wang
摘要
With the rise of Machine Learning as a Service (MLaaS) platforms, safeguarding the intellectual property of deep learning models is becoming paramount. Among various protective measures, trigger set watermarking has emerged as a flexible and effective strategy for preventing unauthorized model distribution. However, this paper identifies an inherent flaw in the current paradigm of trigger set watermarking: evasion adversaries can readily exploit the shortcuts created by models memorizing watermark samples that deviate from the main task distribution, significantly impairing their generalization in adversarial settings. To counteract this, we leverage diffusion models to synthesize unrestricted adversarial examples as trigger sets. By learning the model to accurately recognize them, unique watermark behaviors are promoted through knowledge injection rather than error memorization, thus avoiding exploitable shortcuts. Furthermore, we uncover that the resistance of current trigger set watermarking against removal attacks primarily relies on significantly damaging the decision boundaries during embedding, intertwining unremovability with adverse impacts. By optimizing the knowledge transfer properties of protected models, our approach conveys watermark behaviors to extraction surrogates without aggressively decision boundary perturbation. Experimental results on CIFAR-10/100 and Imagenette datasets demonstrate the effectiveness of our method, showing not only improved robustness against evasion adversaries but also superior resistance to watermark removal attacks compared to state-of-the-art solutions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Revisiting Data Auditing in Large Vision-Language ModelsHongyu Zhu, Sichu Liang, Wenwen Wang, Boheng Li 等ACM MM 2025 · 被引用 3 次
- PATFinger: Prompt-Adapted Transferable Fingerprinting against Unauthorized Multimodal Dataset UsageWenyi Zhang, Ju Jia, Xiaojun Jia, Yihao Huang 等SIGIR 2025 · 被引用 3 次
- AgentMark: Utility-Preserving Behavioral Watermarking for AgentsKaibo Huang, Jin Tan, Yukun Wei, Wanling Li 等ACL 2026 · 被引用 2 次
- Evading Data Provenance in Deep Neural NetworksHongyu Zhu, Sichu Liang, Wenwen Wang, Zhuomeng Zhang 等ICCV 2025
它引用的顶会 Paper49
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
相关 Paper
- Safe and Robust Watermark Injection with a Single OoD ImageShuyang Yu, Junyuan Hong, Haobo Zhang, Haotao Wang 等ICLR 2024 · 被引用 4 次
- Deep Neural Network Watermarking against Model Extraction AttackJingxuan Tan, Nan Zhong, Zhenxing Qian, Xinpeng Zhang 等ACM MM 2023 · 被引用 34 次
- Dataset Reduction and Watermark Removal via Self-supervised Learning for Model Extraction AttackHao Luan, Xue Tan, Zhiheng Li, Jun Dai 等NDSS 2026 · 被引用 1 次
- Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction AttacksYaxin Xiao, Qingqing Ye, Zi Liang, Haoyang Li 等AAAI 2026
- Rethinking the Vulnerability of DNN Watermarking: Are Watermarks Robust against Naturalness-aware Perturbations?Run Wang, Haoxuan Li, Lingzhou Mu, Jixing Ren 等ACM MM 2022 · 被引用 9 次
