Exploring Query Efficient Data Generation Towards Data-Free Model Stealing in Hard Label Setting
Gaozheng Pei, Shaojie Lyu, Ke Ma, Pinci Yang, Qianqian Xu, Yingfei Sun
Abstract
Data-free model stealing involves replicating the functionality of a target model into a substitute model without accessing the target model's structure, parameters, or training data. The adversary can only access the target model's predictions for generated samples. Once the substitute model closely approximates the behavior of the target model, attackers can exploit its white-box characteristics for subsequent malicious activities, such as adversarial attacks. Existing methods within cooperative game frameworks often produce samples with high confidence for the prediction of the substitute model, which makes it difficult for the substitute model to replicate the behavior of the target model. This paper presents a new data-free model stealing approach called Query Efficient Data Generation (QEDG). We introduce two distinct loss functions to ensure the generation of sufficient samples that closely and uniformly align with the target model's decision boundary across multiple classes. Building on the limitation of current methods, which typically yield only one piece of supervised information per query, we propose the queryfree sample augmentation that enables the acquisition of additional supervised information without increasing the number of queries. Motivated by theoretical analysis, we adopt the consistency rate metric, which more accurately evaluates the similarity between the substitute and target models. We conducted extensive experiments to verify the effectiveness of our proposed method, which achieved better performance with fewer queries compared to the state-of-the-art methods on the real MLaaS scenario and five datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2ced1ed-b4ea-48dd-87f9-d89e5344ab3eCited by top-tier papers2
- Divide and Conquer: Heterogeneous Noise Integration for Diffusion-based Adversarial PurificationGaozheng Pei, Shaojie Lyu, Gong Chen, Ke Ma et al.CVPR 2025
- Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMsHaoming Yang, Ke Ma, Xiaojun Jia, Yingfei Sun et al.ICML 2025
Builds on23
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Admix: Enhancing the Transferability of Adversarial AttacksXiaosen Wang, Xuanran He, Jingdong Wang, Kun HeICCV 2021 · 282 citations
- ActiveThief: Model Extraction Using Active Learning and Unannotated Public DataSoham Pal, Yash Gupta, Aditya Shukla, Aditya Kanade et al.AAAI 2020 · 164 citations
- LAS-AT: Adversarial Training with Learnable Attack StrategyXiaojun Jia, Yong Zhang, Baoyuan Wu, Ke Ma et al.CVPR 2022 · 140 citations
- Parallel Rectangle Flip Attack: A Query-based Black-box Attack against Object DetectionSiyuan Liang, Baoyuan Wu, Yanbo Fan, Xingxing Wei et al.ICCV 2021 · 100 citations
Related papers
- Dual Student Networks for Data-Free Model StealingJames Beetham, Navid Kardan, Ajmal Saeed Mian, Mubarak ShahICLR 2023 · 3 citations
- DisGUIDE: Disagreement-Guided Data-Free Model ExtractionJonathan Rosenthal, Eric Enouen, Hung Viet Pham, Lin TanAAAI 2023 · 31 citations
- MAZE: Data-Free Model Stealing Attack Using Zeroth-Order Gradient EstimationSanjay Kariyappa, Atul Prakash, Moinuddin K. QureshiCVPR 2021
- Towards Data-Free Model Stealing in a Hard Label SettingSunandini Sanyal, Sravanti Addepalli, R. Venkatesh BabuCVPR 2022 · 76 citations
- Data-Free Hard-Label Robustness Stealing AttackXiaojian Yuan, Kejiang Chen, Wen Huang, Jie Zhang et al.AAAI 2024 · 11 citations
