Defense against Model Extraction Attack by Bayesian Active Watermarking
Zhenyi Wang, Yihan Wu, Heng Huang
Abstract
Model extraction is to obtain a cloned model that replicates the functionality of a black-box victim model solely through query-based access. Present defense strategies exhibit shortcomings, manifesting as: (1) computational or memory inefficiencies during deployment; or (2) dependence on expensive defensive training methods that mandate the re-training of the victim model; or (3) watermarking-based methods only passively detect model theft without actively preventing model extraction. To address these limitations, we introduce an innovative Bayesian active watermarking technique to fine-tune the victim model and learn the watermark posterior distribution conditioned on input data. The fine-tuning process aims to maximize the log-likelihood on watermarked in-distribution training data for preserving model utility while simultaneously maximizing the change of model's outputs on watermarked out-of-distribution data, thereby achieving effective defense. During deployment, a watermark is randomly sampled from the estimated watermark posterior. This watermark is then added to the input query, and the victim model returns the prediction based on the watermarked input query to users. This proactive defense approach requires only slight fine-tuning of the victim model without the need of full re-training and demonstrates high efficiency in terms of memory and computation during deployment. Rigorous theoretical analysis and comprehensive experimental results demonstrate the efficacy of our proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08739b33-06ba-4a69-93ca-e265b2f93a89Cited by top-tier papers3
- LoRA Recycle: Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAsZixuan Hu, Yongxian Wei, Li Shen, Chun Yuan et al.CVPR 2025
- Dynamic Neural Fortresses: An Adaptive Shield for Model Extraction DefenseSiyu Luan, Zhenyi Wang, Li Shen, Zonghua Gu et al.ICLR 2025
- Adaptive Bayesian Early-Exit Networks for Efficient Non-Transferable LearningSiyu Luan, Yan Li, Zhong Chen, Zhenyi WangCVPR 2026
Builds on29
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas et al.USENIX Security 2018 · 832 citations
- Stealing Hyperparameters in Machine LearningBinghui Wang, Neil Zhenqiang GongS&P 2018 · 504 citations
- Entangled Watermarks as a Defense against Model ExtractionHengrui Jia, Christopher A. Choquette-Choo, Varun Chandrasekaran, Nicolas PapernotUSENIX Security 2021 · 287 citations
Related papers
- Margin-based Neural Network WatermarkingByungjoo Kim, Suyoung Lee, Seanie Lee, Sooel Son et al.ICML 2023 · 21 citations
- DAWN: Dynamic Adversarial Watermarking of Neural NetworksSebastian Szyller, Buse Gul Atli, Samuel Marchal, N. AsokanACM MM 2021 · 133 citations
- Dataset Reduction and Watermark Removal via Self-supervised Learning for Model Extraction AttackHao Luan, Xue Tan, Zhiheng Li, Jun Dai et al.NDSS 2026 · 1 citation
- Thieves on Sesame Street! Model Extraction of BERT-based APIsKalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot et al.ICLR 2020 · 244 citations
- Defending against Data-Free Model Extraction by Distributionally Robust Defensive TrainingZhenyi Wang, Li Shen, Tongliang Liu, Tiehang Duan et al.NeurIPS 2023 · 26 citations
