USENIX Security2018Top-tier venue
With Great Training Comes Great Vulnerability: Practical Attacks against Transfer Learning
Bolun Wang, Yuanshun Yao, Bimal Viswanath, Haitao Zheng, Ben Y. Zhao
Abstract
Transfer learning is a powerful approach that allows users to quickly build accurate deep-learning (Student) models by "learning" from centralized (Teacher) models pretrained with large datasets, e.g. Google's In-ceptionV3. We hypothesize that the centralization of model training increases their vulnerability to misclassification attacks leveraging knowledge of publicly accessible Teacher models. In this paper, we describe our efforts to understand and experimentally validate such attacks in the context of image recognition. We identify techniques that allow attackers to associate Student models with their Teacher counterparts, and launch highly effective misclassification attacks on black-box Student models. We validate this on widely used Teacher models in the wild. Finally, we propose and evaluate multiple approaches for defense, including a neuron-distance technique that successfully defends against these attacks while also obfuscates the link between Teacher and Student models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80d8952a-666f-476b-bd8d-c53dcc91ed8eCited by top-tier papers27
- Latent Backdoor Attacks on Deep Neural NetworksYuanshun Yao, Huiying Li, Haitao Zheng, Ben Y. ZhaoCCS 2019 · 465 citations
- Terminal Brain Damage: Exposing the Graceless Degradation in Deep Neural Networks Under Hardware Fault AttacksSanghyun Hong, Pietro Frigo, Yigitcan Kaya, Cristiano Giuffrida et al.USENIX Security 2019 · 255 citations
- Backdoor Scanning for Deep Neural Networks through K-Arm OptimizationGuangyu Shen, Yingqi Liu, Guanhong Tao, Shengwei An et al.ICML 2021 · 137 citations
- T-Miner: A Generative Approach to Defend Against Trojan Attacks on DNN-based Text ClassificationAhmadreza Azizi, Ibrahim Asadullah Tahmid, Asim Waheed, Neal Mangaokar et al.USENIX Security 2021 · 98 citations
- Poisoning the Unlabeled Dataset of Semi-Supervised LearningNicholas CarliniUSENIX Security 2021 · 80 citations
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 1,765 citations
Related papers
- A Target-Agnostic Attack on Deep Models: Exploiting Security Vulnerabilities of Transfer LearningShahbaz Rezaei, Xin LiuICLR 2020 · 49 citations
- Perturbing Across the Feature Hierarchy to Improve Standard and Strict Blackbox Attack TransferabilityNathan Inkawhich, Kevin J. Liang, Binghui Wang, Matthew Inkawhich et al.NeurIPS 2020 · 105 citations
- Transferable Perturbations of Deep Feature DistributionsNathan Inkawhich, Kevin J. Liang, Lawrence Carin, Yiran ChenICLR 2020 · 100 citations
- AdvFlow: Inconspicuous Black-box Adversarial Attacks using Normalizing FlowsHadi Mohaghegh Dolatabadi, Sarah M. Erfani, Christopher LeckieNeurIPS 2020 · 75 citations
- Black-Box Adversarial Attack with Transferable Model-based EmbeddingZhichao Huang, Tong ZhangICLR 2020 · 131 citations
