D-DAE: Defense-Penetrating Model Extraction Attacks
Yanjiao Chen, Rui Guan, Xueluan Gong, Jianshuo Dong, Meng Xue
Abstract
Recent studies show that machine learning models are vulnerable to model extraction attacks, where the adversary builds a substitute model that achieves almost the same performance of a black-box victim model simply via querying the victim model. To defend against such attacks, a series of methods have been proposed to disrupt the query results before returning them to potential attackers, greatly degrading the performance of existing model extraction attacks.In this paper, we make the first attempt to develop a defense-penetrating model extraction attack framework, named D-DAE, which aims to break disruption-based defenses. The linchpins of D-DAE are the design of two modules, i.e., disruption detection and disruption recovery, which can be integrated with generic model extraction attacks. More specifically, after obtaining query results from the victim model, the disruption detection module infers the defense mechanism adopted by the defender. We design a meta-learning-based disruption detection algorithm for learning the fundamental differences between the distributions of disrupted and undisrupted query results. The algorithm features a good generalization property even if we have no access to the original training dataset of the victim model. Given the detected defense mechanism, the disruption recovery module tries to restore a clean query result from the disrupted query result with well-designed generative models. Our extensive evaluations on MNIST, FashionMNIST, CIFAR-10, GTSRB, and ImageNette datasets demonstrate that D-DAE can enhance the substitute model accuracy of the existing model extraction attacks by as much as 82.24% in the face of 4 state-of-the-art defenses and combinations of multiple defenses. We also verify the effectiveness of D-DAE in penetrating unknown defenses in real-world APIs hosted by Microsoft Azure and Face++.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cb948988-9b0e-4c81-bcdb-32ffe4c1b245Cited by top-tier papers7
- PromptCARE: Prompt Copyright Protection by Watermark Injection and VerificationHongwei Yao, Jian Lou, Zhan Qin, Kui RenS&P 2024 · 43 citations
- ModelGuard: Information-Theoretic Defense Against Model Extraction AttacksMinxue Tang, Anna Dai, Louis DiValentin, Aolin Ding et al.USENIX Security 2024 · 28 citations
- Efficient Model Stealing Defense with Noise Transition MatrixDong-Dong Wu, Chilin Fu, Weichang Wu, Wenwen Xia et al.CVPR 2024 · 2 citations
- Dataset Reduction and Watermark Removal via Self-supervised Learning for Model Extraction AttackHao Luan, Xue Tan, Zhiheng Li, Jun Dai et al.NDSS 2026 · 1 citation
- Towards Understanding and Enhancing Security of Proof-of-Training for DNN Model Ownership VerificationYijia Chang, Hanrui Jiang, Chao Lin, Xinyi Huang et al.USENIX Security 2025
Related papers
- Exploring Connections Between Active Learning and Model ExtractionVarun Chandrasekaran, Kamalika Chaudhuri, Irene Giacomelli, Somesh Jha et al.USENIX Security 2020
- SAME: Sample Reconstruction against Model Extraction AttacksYi Xie, Jie Zhang, Shiqian Zhao, Tianwei Zhang et al.AAAI 2024 · 6 citations
- CloudLeak: Large-Scale Deep Learning Models Stealing Through Adversarial ExamplesHonggang Yu, Kaichen Yang, Teng Zhang, Yun-Yun Tsai et al.NDSS 2020
- Beowulf: Mitigating Model Extraction Attacks Via Reshaping Decision RegionsXueluan Gong, Rubin Wei, Ziyao Wang, Yuchen Sun et al.CCS 2024 · 2 citations
- Extracting Robust Models with Uncertain ExamplesGuanlin Li, Guowen Xu, Shangwei Guo, Han Qiu et al.ICLR 2023
