Thieves on Sesame Street! Model Extraction of BERT-based APIs
Kalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot, Mohit Iyyer
摘要
We study the problem of model extraction in natural language processing, in which an adversary with only query access to a victim model attempts to reconstruct a local copy of that model. Assuming that both the adversary and victim model fine-tune a large pretrained language model such as BERT (Devlin et al. 2019), we show that the adversary does not need any real training data to successfully mount the attack. In fact, the attacker need not even use grammatical or semantically meaningful queries: we show that random sequences of words coupled with task-specific heuristics form effective queries for model extraction on a diverse set of NLP tasks, including natural language inference and question answering. Our work thus highlights an exploit only made feasible by the shift towards transfer learning methods within the NLP community: for a query budget of a few hundred dollars, an attacker can extract a model that performs only slightly worse than the victim model. Finally, we study two defense strategies against model extraction---membership classification and API watermarking---which while successful against naive adversaries, are ineffective against more sophisticated ones.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper65
- Adversarial Watermarking Transformer: Towards Tracing Text Provenance with Data HidingSahar Abdelnabi, Mario FritzS&P 2021 · 被引用 210 次
- Deep Neural Network Fingerprinting by Conferrable Adversarial ExamplesNils Lukas, Yuxuan Zhang, Florian KerschbaumICLR 2021 · 被引用 182 次
- Protecting Intellectual Property of Language Generation APIs with Lexical WatermarkXuanli He, Qiongkai Xu, Lingjuan Lyu, Fangzhao Wu 等AAAI 2022 · 被引用 124 次
- SoK: How Robust is Image Classification Deep Neural Network Watermarking?Nils Lukas, Edward Jiang, Xinda Li, Florian KerschbaumS&P 2022 · 被引用 124 次
- Knowledge-Enriched Distributional Model Inversion AttacksSi Chen, Mostafa Kahla, Ruoxi Jia, Guo-Jun QiICCV 2021 · 被引用 124 次
它引用的顶会 Paper7
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
- Machine Learning with Membership Privacy using Adversarial RegularizationMilad Nasr, Reza Shokri, Amir HoumansadrCCS 2018 · 被引用 543 次
相关 Paper
- MeaeQ: Mount Model Extraction Attacks with Efficient QueriesChengwei Dai, Minxuan Lv, Kun Li, Wei ZhouEMNLP 2023 · 被引用 4 次
- Grey-box Extraction of Natural Language ModelsSantiago Zanella-Béguelin, Shruti Tople, Andrew Paverd, Boris KöpfICML 2021 · 被引用 38 次
- Marich: A Query-efficient Distributionally Equivalent Model Extraction AttackPratik Karmakar, Debabrota BasuNeurIPS 2023 · 被引用 11 次
- Fragments to Facts: Partial-Information Fragment Inference from LLMsLucas Rosenblatt, Bin Han, Robert Wolfe, Bill HoweICML 2025
- BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation ModelsKangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo 等ICLR 2022 · 被引用 133 次
