Margin-based Neural Network Watermarking
Byungjoo Kim, Suyoung Lee, Seanie Lee, Sooel Son, Sung Ju Hwang
Abstract
As Machine Learning as a Service (MLaaS) platforms become prevalent, deep neural network (DNN) watermarking techniques are gaining increasing attention, which enables one to verify the ownership of a target DNN model in a black-box scenario. Unfortunately, previous watermarking methods are vulnerable to functionality stealing attacks, thus allowing an adversary to falsely claim the ownership of a DNN model stolen from its original owner. In this work, we propose a novel margin-based DNN watermarking approach that is robust to the functionality stealing attacks based on model extraction and distillation. Specifically, during training, our method maximizes the margins of watermarked samples by using projected gradient ascent on them so that their predicted labels cannot change without compromising the accuracy of the model that the attacker tries to steal. We validate our method on multiple benchmarks and show that our watermarking method successfully defends against model extraction attacks, outperforming relevant baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dcec374d-ba9e-412a-972a-0c8b481b223bCited by top-tier papers7
- Reliable Model Watermarking: Defending against Theft without Compromising on EvasionHongyu Zhu, Sichu Liang, Wentao Hu, Fangqi Li et al.ACM MM 2024 · 14 citations
- Self-Supervised Dataset Distillation for Transfer LearningDong Bok Lee, Seanie Lee, Joonho Ko, Kenji Kawaguchi et al.ICLR 2024 · 9 citations
- Towards the Resistance of Neural Network Fingerprinting to Fine-tuningLing Tang, Yuefeng Chen, Hui Xue', Quanshi ZhangNeurIPS 2025 · 5 citations
- PlugMark: A Plug-In Zero-Watermarking Framework for Diffusion ModelsPengzhen Chen, Yanwei Liu, Xiaoyan Gu, Enci Liu et al.ICCV 2025 · 1 citation
- Dataset Reduction and Watermark Removal via Self-supervised Learning for Model Extraction AttackHao Luan, Xue Tan, Zhiheng Li, Jun Dai et al.NDSS 2026 · 1 citation
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas et al.USENIX Security 2018 · 832 citations
- Entangled Watermarks as a Defense against Model ExtractionHengrui Jia, Christopher A. Choquette-Choo, Varun Chandrasekaran, Nicolas PapernotUSENIX Security 2021 · 287 citations
- Deep Neural Network Fingerprinting by Conferrable Adversarial ExamplesNils Lukas, Yuxuan Zhang, Florian KerschbaumICLR 2021 · 182 citations
Related papers
- DAWN: Dynamic Adversarial Watermarking of Neural NetworksSebastian Szyller, Buse Gul Atli, Samuel Marchal, N. AsokanACM MM 2021 · 133 citations
- MEA-Defender: A Robust Watermark against Model Extraction AttackPeizhuo Lv, Hualong Ma, Kai Chen, Jiachen Zhou et al.S&P 2024 · 22 citations
- DeepTracer: Tracing Stolen Model via Deep Coupled WatermarksYunfei Yang, Xiaojun Chen, Yuexin Xuan, Zhendong Zhao et al.AAAI 2026
- Towards Robust Model Watermark via Reducing Parametric VulnerabilityGuanhao Gan, Yiming Li, Dongxian Wu, Shu-Tao XiaICCV 2023 · 18 citations
- Neural Dehydration: Effective Erasure of Black-box Watermarks from DNNs with Limited DataYifan Lu, Wenxuan Li, Mi Zhang, Xudong Pan et al.CCS 2024 · 2 citations
