Deep Neural Network Watermarking against Model Extraction Attack
Jingxuan Tan, Nan Zhong, Zhenxing Qian, Xinpeng Zhang, Sheng Li
Abstract
Deep neural network (DNN) watermarking is an emerging technique to protect the intellectual property of deep learning models. At present, many DNN watermarking algorithms have been proposed to achieve provenance verification by embedding identify information into the internals or prediction behaviors of the host model. However, most methods are vulnerable to model extraction attacks, where attackers collect output labels from the model to train a surrogate or a replica. To address this issue, we present a novel DNN watermarking approach, named SSW, which constructs an adaptive trigger set progressively by optimizing over a pair of symmetric shadow models to enhance the robustness to model extraction. Precisely, we train a positive shadow model supervised by the prediction of the host model to mimic the behaviors of potential surrogate models. Additionally, a negative shadow model is normally trained to imitate irrelevant independent models. Using this pair of shadow models as a reference, we design a strategy to update the trigger samples appropriately such that they tend to persist in the host model and its stolen copies. Moreover, our method could well support two specific embedding schemes: embedding the watermark via fine-tuning or from scratch. Our extensive experimental results on popular datasets demonstrate that our SSW approach outperforms state-of-the-art methods against various model extraction attacks in whether trigger set classification accuracy based or hypothesis test based verification. The results also show that our method is robust to common model modification schemes including fine-tuning and model compression.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2a9431a9-026d-45cb-8514-d89a0c9e9599Cited by top-tier papers7
- Reliable Model Watermarking: Defending against Theft without Compromising on EvasionHongyu Zhu, Sichu Liang, Wentao Hu, Fangqi Li et al.ACM MM 2024 · 14 citations
- MFL-Owner: Ownership Protection for Multi-modal Federated Learning via Orthogonal Transform WatermarkKeke Gai, Dongjue Wang, Jing Yu, Mohan Wang et al.AAAI 2025 · 6 citations
- HoneypotNet: Backdoor Attacks Against Model ExtractionYixu Wang, Tianle Gu, Yan Teng, Yingchun Wang et al.AAAI 2025 · 4 citations
- CREDIT: Certified Ownership Verification of Deep Neural Networks Against Model Extraction AttacksBolin Shen, Zhan Cheng, Neil Gong, Fan Yao et al.ICML 2026 · 3 citations
- CodeGenGuard: A Watermark for Code Generation ModelsBorui Yang, Mingxuan Ma, Liyao Xiang, Nan Chen et al.ICLR 2026
Related papers
- SoK: How Robust is Image Classification Deep Neural Network Watermarking?Nils Lukas, Edward Jiang, Xinda Li, Florian KerschbaumS&P 2022 · 124 citations
- MEA-Defender: A Robust Watermark against Model Extraction AttackPeizhuo Lv, Hualong Ma, Kai Chen, Jiachen Zhou et al.S&P 2024 · 22 citations
- Identification for Deep Neural Network: Simply Adjusting Few Weights!Yingjie Lao, Peng Yang, Weijie Zhao, Ping LiICDE 2022 · 19 citations
- Watermarking Deep Neural Networks with Greedy ResidualsHanwen Liu, Zhenyu Weng, Yuesheng ZhuICML 2021 · 69 citations
- DAWN: Dynamic Adversarial Watermarking of Neural NetworksSebastian Szyller, Buse Gul Atli, Samuel Marchal, N. AsokanACM MM 2021 · 133 citations
