MAZE: Data-Free Model Stealing Attack Using Zeroth-Order Gradient Estimation
Sanjay Kariyappa, Atul Prakash, Moinuddin K. Qureshi
Abstract
High quality Machine Learning (ML) models are often considered valuable intellectual property by companies. Model Stealing (MS) attacks allow an adversary with blackbox access to a ML model to replicate its functionality by training a clone model using the predictions of the target model for different inputs. However, best available existing MS attacks fail to produce a high-accuracy clone without access to the target dataset or a representative dataset necessary to query the target model. In this paper, we show that preventing access to the target dataset is not an adequate defense to protect a model. We propose MAZE -a data-free model stealing attack using zeroth-order gradient estimation that produces high-accuracy clones. In contrast to prior works, MAZE uses only synthetic data created using a generative model to perform MS. Our evaluation with four image classification models shows that MAZE provides a normalized clone accuracy in the range of 0.90⇥ to 0.99⇥, and outperforms even the recent attacks that rely on partial data (JBDA, clone accuracy 0.13⇥ to 0.69⇥) and on surrogate data (KnockoffNets, clone accuracy 0.52⇥ to 0.97⇥). We also study an extension of MAZE in the partial-data setting, and develop MAZE-PD, which generates synthetic data closer to the target distribution. MAZE-PD further improves the clone accuracy (0.97⇥ to 1.0⇥) and reduces the query budget required for the attack by 2⇥-24⇥.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 371d92dd-9cdd-45f8-b088-9b36ba74f3daCited by top-tier papers47
- Hermes Attack: Steal DNN Models with Lossless Inference AccuracyYuankun Zhu, Yueqiang Cheng, Husheng Zhou, Yantao LuUSENIX Security 2021 · 119 citations
- Towards Data-Free Model Stealing in a Hard Label SettingSunandini Sanyal, Sravanti Addepalli, R. Venkatesh BabuCVPR 2022 · 76 citations
- DisGUIDE: Disagreement-Guided Data-Free Model ExtractionJonathan Rosenthal, Eric Enouen, Hung Viet Pham, Lin TanAAAI 2023 · 31 citations
- Defending against Data-Free Model Extraction by Distributionally Robust Defensive TrainingZhenyi Wang, Li Shen, Tongliang Liu, Tiehang Duan et al.NeurIPS 2023 · 26 citations
- StolenEncoder: Stealing Pre-trained Encoders in Self-supervised LearningYupei Liu, Jinyuan Jia, Hongbin Liu, Neil Zhenqiang GongCCS 2022 · 23 citations
Builds on7
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Stealing Hyperparameters in Machine LearningBinghui Wang, Neil Zhenqiang GongS&P 2018 · 504 citations
- Prediction Poisoning: Towards Defenses Against DNN Model Stealing AttacksTribhuvanesh Orekondy, Bernt Schiele, Mario FritzICLR 2020 · 194 citations
- Defending Against Model Stealing Attacks With Adaptive MisinformationSanjay Kariyappa, Moinuddin K. QureshiCVPR 2020
Related papers
- Dual Student Networks for Data-Free Model StealingJames Beetham, Navid Kardan, Ajmal Saeed Mian, Mubarak ShahICLR 2023 · 3 citations
- Data-Free Hard-Label Robustness Stealing AttackXiaojian Yuan, Kejiang Chen, Wen Huang, Jie Zhang et al.AAAI 2024 · 11 citations
- Exploring Query Efficient Data Generation Towards Data-Free Model Stealing in Hard Label SettingGaozheng Pei, Shaojie Lyu, Ke Ma, Pinci Yang et al.AAAI 2025 · 2 citations
- Data-Free Model ExtractionJean-Baptiste Truong, Pratyush Maini, Robert J. Walls, Nicolas PapernotCVPR 2021
- Protecting DNNs from Theft using an Ensemble of Diverse ModelsSanjay Kariyappa, Atul Prakash, Moinuddin K. QureshiICLR 2021 · 33 citations
