Grey-box Extraction of Natural Language Models
Santiago Zanella-Béguelin, Shruti Tople, Andrew Paverd, Boris Köpf
摘要
Model extraction attacks attempt to replicate a target machine learning model from predictions obtained by querying its inference API. Most existing attacks on Deep Neural Networks achieve this by supervised training of the copy using the victim's predictions. An emerging class of attacks exploit algebraic properties of DNNs to obtain high-fidelity copies using orders of magnitude fewer queries than the prior state-of-the-art. So far, such powerful attacks have been limited to networks with few hidden layers and ReLU activations. In this paper we present algebraic attacks on large-scale natural language models in a grey-box setting, targeting models with a pre-trained (public) encoder followed by a single (private) classification layer. Our key observation is that a small set of arbitrary embedding vectors is likely to form a basis of the classification layer's input space, which a grey-box adversary can compute. We show how to use this information to solve an equation system that determines the classification layer from the corresponding probability outputs. We evaluate the effectiveness of our attacks on different sizes of transformer models and downstream tasks. Our key findings are that (i) with frozen base layers, high-fidelity extraction is possible with a number of queries that is as small as twice the input dimension of the last layer. This is true even for queries that are entirely in-distribution, making extraction attacks indistinguishable from legitimate use; (ii) with fine-tuned base layers, the effectiveness of algebraic attacks decreases with the learning rate, showing that fine-tuning is not only beneficial for accuracy but also indispensable for model confidentiality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Stealing part of a production language modelNicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke 等ICML 2024 · 被引用 157 次
- Fingerprinting Deep Neural Networks Globally via Universal Adversarial PerturbationsZirui Peng, Shaofeng Li, Guoxing Chen, Cheng Zhang 等CVPR 2022 · 被引用 66 次
- Increasing the Cost of Model Extraction with Calibrated Proof of WorkAdam Dziedzic, Muhammad Ahmad Kaleem, Yu Shen Lu, Nicolas PapernotICLR 2022 · 被引用 37 次
- Defending against Data-Free Model Extraction by Distributionally Robust Defensive TrainingZhenyi Wang, Li Shen, Tongliang Liu, Tiehang Duan 等NeurIPS 2023 · 被引用 26 次
- StolenEncoder: Stealing Pre-trained Encoders in Self-supervised LearningYupei Liu, Jinyuan Jia, Hongbin Liu, Neil Zhenqiang GongCCS 2022 · 被引用 23 次
它引用的顶会 Paper9
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
- Thieves on Sesame Street! Model Extraction of BERT-based APIsKalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot 等ICLR 2020 · 被引用 244 次
- Prediction Poisoning: Towards Defenses Against DNN Model Stealing AttacksTribhuvanesh Orekondy, Bernt Schiele, Mario FritzICLR 2020 · 被引用 194 次
- ActiveThief: Model Extraction Using Active Learning and Unannotated Public DataSoham Pal, Yash Gupta, Aditya Shukla, Aditya Kanade 等AAAI 2020 · 被引用 164 次
- Reverse-engineering deep ReLU networksDavid Rolnick, Konrad P. KordingICML 2020 · 被引用 121 次
相关 Paper
- High Accuracy and High Fidelity Extraction of Neural NetworksMatthew Jagielski, Nicholas Carlini, David Berthelot, Alex Kurakin 等USENIX Security 2020
- Cryptanalytic Extraction of Neural Network ModelsNicholas Carlini, Matthew Jagielski, Ilya MironovCRYPTO 2020 · 被引用 109 次
- Teach LLMs to Phish: Stealing Private Information from Language ModelsAshwinee Panda, Christopher A. Choquette-Choo, Zhengming Zhang, Yaoqing Yang 等ICLR 2024 · 被引用 41 次
- Is the Hard-Label Cryptanalytic Model Extraction Really Polynomial?Akira Ito, Takayuki Miura, Yosuke TodoCRYPTO 2026
- Combing for Credentials: Active Pattern Extraction from Smart ReplyBargav Jayaraman, Esha Ghosh, Melissa Chase, Sambuddha Roy 等S&P 2024 · 被引用 11 次
