Defending against Model Stealing via Verifying Embedded External Features
Yiming Li, Linghui Zhu, Xiaojun Jia, Yong Jiang, Shu-Tao Xia, Xiaochun Cao
Abstract
Obtaining a well-trained model involves expensive data collection and training procedures, therefore the model is a valuable intellectual property. Recent studies revealed that adversaries can 'steal' deployed models even when they have no training samples and can not get access to the model parameters or structures. Currently, there were some defense methods to alleviate this threat, mostly by increasing the cost of model stealing. In this paper, we explore the defense from another angle by verifying whether a suspicious model contains the knowledge of defender-specified external features. Specifically, we embed the external features by tempering a few training samples with style transfer. We then train a meta-classifier to determine whether a model is stolen from the victim. This approach is inspired by the understanding that the stolen models should contain the knowledge of features learned by the victim model. We examine our method on both CIFAR-10 and ImageNet datasets. Experimental results demonstrate that our method is effective in detecting different types of model stealing simultaneously, even if the stolen model is obtained via a multi-stage stealing process. The codes for reproducing main results are available at Github ( https://github.com/zlh-thu/StealingVerification ). Recently, researchers found that adversaries can 'steal' the (deployed) victim model even when they have no training samples and can not get access to the model parameters or structures (Tramèr et al. 2016; Orekondy, Schiele, and Fritz 2019; Chandrasekaran et al. 2020) . For example, the adversaries may use the victim model to label an unlabeled dataset, based on which to train the stolen model. This threat
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a0409c7-ba9e-451b-9200-30dac500652bCited by top-tier papers26
- Untargeted Backdoor Watermark: Towards Harmless and Stealthy Dataset Copyright ProtectionYiming Li, Yang Bai, Yong Jiang, Yong Yang et al.NeurIPS 2022 · 161 citations
- Are You Stealing My Model? Sample Correlation for Fingerprinting Deep Neural NetworksJiyang Guan, Jian Liang, Ran HeNeurIPS 2022 · 57 citations
- IBD-PSC: Input-level Backdoor Detection via Parameter-oriented Scaling ConsistencyLinshan Hou, Ruili Feng, Zhongyun Hua, Wei Luo et al.ICML 2024 · 52 citations
- HuRef: HUman-REadable Fingerprint for Large Language ModelsBoyi Zeng, Lizheng Wang, Yuncong Hu, Yi Xu et al.NeurIPS 2024 · 48 citations
- PromptCARE: Prompt Copyright Protection by Watermark Injection and VerificationHongwei Yao, Jian Lou, Zhan Qin, Kui RenS&P 2024 · 43 citations
Builds on20
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas et al.USENIX Security 2018 · 832 citations
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 601 citations
- Property Inference Attacks on Fully Connected Neural Networks using Permutation Invariant RepresentationsKaran Ganju, Qi Wang, Wei Yang, Carl A. Gunter et al.CCS 2018 · 574 citations
Related papers
- Dataset Inference: Ownership Resolution in Machine LearningPratyush Maini, Mohammad Yaghini, Nicolas PapernotICLR 2021 · 155 citations
- Entangled Watermarks as a Defense against Model ExtractionHengrui Jia, Christopher A. Choquette-Choo, Varun Chandrasekaran, Nicolas PapernotUSENIX Security 2021 · 287 citations
- Fingerprinting Deep Neural Networks Globally via Universal Adversarial PerturbationsZirui Peng, Shaofeng Li, Guoxing Chen, Cheng Zhang et al.CVPR 2022 · 66 citations
- Dataset Inference for Self-Supervised ModelsAdam Dziedzic, Haonan Duan, Muhammad Ahmad Kaleem, Nikita Dhawan et al.NeurIPS 2022 · 59 citations
- Can't Steal? Cont-Steal! Contrastive Stealing Attacks Against Image EncodersZeyang Sha, Xinlei He, Ning Yu, Michael Backes et al.CVPR 2023
