ActiveThief: Model Extraction Using Active Learning and Unannotated Public Data
Soham Pal, Yash Gupta, Aditya Shukla, Aditya Kanade, Shirish K. Shevade, Vinod Ganapathy
Abstract
Machine learning models are increasingly being deployed in practice. Machine Learning as a Service (MLaaS) providers expose such models to queries by third-party developers through application programming interfaces (APIs). Prior work has developed model extraction attacks, in which an attacker extracts an approximation of an MLaaS model by making black-box queries to it. We design ActiveThief – a model extraction framework for deep neural networks that makes use of active learning techniques and unannotated public datasets to perform model extraction. It does not expect strong domain knowledge or access to annotated data on the part of the attacker. We demonstrate that (1) it is possible to use ActiveThief to extract deep classifiers trained on a variety of datasets from image and text domains, while querying the model with as few as 10-30% of samples from public datasets, (2) the resulting model exhibits a higher transferability success rate of adversarial examples than prior work, and (3) the attack evades detection by the state-of-the-art model extraction detection method, PRADA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4121a8c1-1d2c-45a4-bd25-30c6971e4022Cited by top-tier papers31
- Students Parrot Their Teachers: Membership Inference on Model DistillationMatthew Jagielski, Milad Nasr, Katherine Lee, Christopher A. Choquette-Choo et al.NeurIPS 2023 · 53 citations
- Privacy Side Channels in Machine Learning SystemsEdoardo Debenedetti, Giorgio Severi, Milad Nasr, Christopher A. Choquette-Choo et al.USENIX Security 2024 · 52 citations
- How to Steer Your Adversary: Targeted and Efficient Model Stealing Defenses with Gradient RedirectionMantas Mazeika, Bo Li, David A. ForsythICML 2022 · 39 citations
- Grey-box Extraction of Natural Language ModelsSantiago Zanella-Béguelin, Shruti Tople, Andrew Paverd, Boris KöpfICML 2021 · 38 citations
- Increasing the Cost of Model Extraction with Calibrated Proof of WorkAdam Dziedzic, Muhammad Ahmad Kaleem, Yu Shen Lu, Nicolas PapernotICLR 2022 · 37 citations
Builds on2
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Exploring Connections Between Active Learning and Model ExtractionVarun Chandrasekaran, Kamalika Chaudhuri, Irene Giacomelli, Somesh Jha et al.USENIX Security 2020
Related papers
- CloudLeak: Large-Scale Deep Learning Models Stealing Through Adversarial ExamplesHonggang Yu, Kaichen Yang, Teng Zhang, Yun-Yun Tsai et al.NDSS 2020
- Practical and Efficient Model Extraction of Sentiment Analysis APIsWeibin Wu, Jianping Zhang, Victor Junqiu Wei, Xixian Chen et al.ICSE 2023 · 10 citations
- Extracting Robust Models with Uncertain ExamplesGuanlin Li, Guowen Xu, Shangwei Guo, Han Qiu et al.ICLR 2023
- Deep Neural Network Fingerprinting by Conferrable Adversarial ExamplesNils Lukas, Yuxuan Zhang, Florian KerschbaumICLR 2021 · 182 citations
- On the Difficulty of Defending Self-Supervised Learning against Model ExtractionAdam Dziedzic, Nikita Dhawan, Muhammad Ahmad Kaleem, Jonas Guan et al.ICML 2022 · 34 citations
