How to Steer Your Adversary: Targeted and Efficient Model Stealing Defenses with Gradient Redirection
Mantas Mazeika, Bo Li, David A. Forsyth
摘要
Model stealing attacks present a dilemma for public machine learning APIs. To protect financial investments, companies may be forced to withhold important information about their models that could facilitate theft, including uncertainty estimates and prediction explanations. This compromise is harmful not only to users but also to external transparency. Model stealing defenses seek to resolve this dilemma by making models harder to steal while preserving utility for benign users. However, existing defenses have poor performance in practice, either requiring enormous computational overheads or severe utility trade-offs. To meet these challenges, we present a new approach to model stealing defenses called gradient redirection. At the core of our approach is a provably optimal, efficient algorithm for steering an adversary's training updates in a targeted manner. Combined with improvements to surrogate networks and a novel coordinated defense strategy, our gradient redirection defense, called GRAD, achieves small utility trade-offs and low computational overhead, outperforming the best prior defenses. Moreover, we demonstrate how gradient redirection enables reprogramming the adversary with arbitrary behavior, which we hope will foster work on new avenues of defense.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- ModelGuard: Information-Theoretic Defense Against Model Extraction AttacksMinxue Tang, Anna Dai, Louis DiValentin, Aolin Ding 等USENIX Security 2024 · 被引用 28 次
- Defending against Data-Free Model Extraction by Distributionally Robust Defensive TrainingZhenyi Wang, Li Shen, Tongliang Liu, Tiehang Duan 等NeurIPS 2023 · 被引用 26 次
- SoK: All You Need to Know About On-Device ML Model Extraction - The Gap Between Research and PracticeTushar Nayan, Qiming Guo, Mohammed Alduniawi, Marcus Botacin 等USENIX Security 2024 · 被引用 20 次
- Medical Multimodal Model Stealing Attacks via Adversarial Domain AlignmentYaling Shen, Zhixiong Zhuang, Kun Yuan, Maria-Irina Nicolae 等AAAI 2025 · 被引用 12 次
- ModelObfuscator: Obfuscating Model Information to Protect Deployed ML-Based SystemsMingyi Zhou, Xiang Gao, Jing Wu, John C. Grundy 等ISSTA 2023 · 被引用 11 次
它引用的顶会 Paper7
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas 等USENIX Security 2018 · 被引用 832 次
- Thieves on Sesame Street! Model Extraction of BERT-based APIsKalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot 等ICLR 2020 · 被引用 244 次
- Prediction Poisoning: Towards Defenses Against DNN Model Stealing AttacksTribhuvanesh Orekondy, Bernt Schiele, Mario FritzICLR 2020 · 被引用 194 次
- ActiveThief: Model Extraction Using Active Learning and Unannotated Public DataSoham Pal, Yash Gupta, Aditya Shukla, Aditya Kanade 等AAAI 2020 · 被引用 164 次
相关 Paper
- Increasing the Cost of Model Extraction with Calibrated Proof of WorkAdam Dziedzic, Muhammad Ahmad Kaleem, Yu Shen Lu, Nicolas PapernotICLR 2022 · 被引用 37 次
- Efficient Model Stealing Defense with Noise Transition MatrixDong-Dong Wu, Chilin Fu, Weichang Wu, Wenwen Xia 等CVPR 2024 · 被引用 2 次
- Isolation and Induction: Training Robust Deep Neural Networks against Model Stealing AttacksJun Guo, Xingyu Zheng, Aishan Liu, Siyuan Liang 等ACM MM 2023 · 被引用 8 次
- Defending Against Model Stealing Attacks With Adaptive MisinformationSanjay Kariyappa, Moinuddin K. QureshiCVPR 2020
- Bucks for Buckets (B4B): Active Defenses Against Stealing EncodersJan Dubinski, Stanislaw Pawlak, Franziska Boenisch, Tomasz Trzcinski 等NeurIPS 2023 · 被引用 12 次
