Protecting DNNs from Theft using an Ensemble of Diverse Models
Sanjay Kariyappa, Atul Prakash, Moinuddin K. Qureshi
Abstract
Several recent works have demonstrated highly effective model stealing (MS) attacks on Deep Neural Networks (DNNs) in black-box settings, even when the training data is unavailable. These attacks typically use some form of Out of Distribution (OOD) data to query the target model and use the predictions obtained to train a clone model. Such a clone model learns to approximate the decision boundary of the target model, achieving high accuracy on in-distribution examples. We propose Ensemble of Diverse Models (EDM) to defend against such MS attacks. EDM is made up of models that are trained to produce dissimilar predictions for OOD inputs. By using a different member of the ensemble to service different queries, our defense produces predictions that are highly discontinuous in the input space for the adversary's OOD queries. Such discontinuities cause the clone model trained on these predictions to have poor generalization on in-distribution examples. Our evaluations on several image classification tasks demonstrate that EDM defense can severely degrade the accuracy of clone models (up to ). Our defense has minimal impact on the target accuracy, negligible computational costs during inference, and is compatible with existing defenses for MS attacks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1de6ece7-894f-4cdf-812e-58e12634c370Cited by top-tier papers9
- Towards Data-Free Model Stealing in a Hard Label SettingSunandini Sanyal, Sravanti Addepalli, R. Venkatesh BabuCVPR 2022 · 76 citations
- Increasing the Cost of Model Extraction with Calibrated Proof of WorkAdam Dziedzic, Muhammad Ahmad Kaleem, Yu Shen Lu, Nicolas PapernotICLR 2022 · 37 citations
- Defending against Data-Free Model Extraction by Distributionally Robust Defensive TrainingZhenyi Wang, Li Shen, Tongliang Liu, Tiehang Duan et al.NeurIPS 2023 · 26 citations
- Defense against Model Extraction Attack by Bayesian Active WatermarkingZhenyi Wang, Yihan Wu, Heng HuangICML 2024 · 10 citations
- Isolation and Induction: Training Robust Deep Neural Networks against Model Stealing AttacksJun Guo, Xingyu Zheng, Aishan Liu, Siyuan Liang et al.ACM MM 2023 · 8 citations
Related papers
- Defending Against Model Stealing Attacks With Adaptive MisinformationSanjay Kariyappa, Moinuddin K. QureshiCVPR 2020
- Prediction Poisoning: Towards Defenses Against DNN Model Stealing AttacksTribhuvanesh Orekondy, Bernt Schiele, Mario FritzICLR 2020 · 194 citations
- Simulating Unknown Target Models for Query-Efficient Black-Box AttacksChen Ma, Li Chen, Jun-Hai YongCVPR 2021
- Delving into Data: Effectively Substitute Training for Black-box AttackWenxuan Wang, Bangjie Yin, Taiping Yao, Li Zhang et al.CVPR 2021
- MAZE: Data-Free Model Stealing Attack Using Zeroth-Order Gradient EstimationSanjay Kariyappa, Atul Prakash, Moinuddin K. QureshiCVPR 2021
