Simulating Unknown Target Models for Query-Efficient Black-Box Attacks
Chen Ma, Li Chen, Jun-Hai Yong
摘要
Many adversarial attacks have been proposed to investigate the security issues of deep neural networks. In the black-box setting, current model stealing attacks train a substitute model to counterfeit the functionality of the target model. However, the training requires querying the target model. Consequently, the query complexity remains high, and such attacks can be defended easily. This study aims to train a generalized substitute model called "Simulator", which can mimic the functionality of any unknown target model. To this end, we build the training data with the form of multiple tasks by collecting query sequences generated during the attacks of various existing networks. The learning process uses a mean square error-based knowledgedistillation loss in the meta-learning to minimize the difference between the Simulator and the sampled networks. The meta-gradients of this loss are then computed and accumulated from multiple tasks to update the Simulator and subsequently improve generalization. When attacking a target model that is unseen in training, the trained Simulator can accurately simulate its functionality using its limited feedback. As a result, a large fraction of queries can be transferred to the Simulator, thereby reducing query complexity. Results of the comprehensive experiments conducted using the CIFAR-10, CIFAR-100, and TinyImageNet datasets demonstrate that the proposed approach reduces query complexity by several orders of magnitude compared to the baseline method. The implementation source code is released online 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Blackbox Attacks via Surrogate Ensemble SearchZikui Cai, Chengyu Song, Srikanth V. Krishnamurthy, Amit Roy-Chowdhury 等NeurIPS 2022 · 被引用 32 次
- Finding Optimal Tangent Points for Reducing Distortions of Hard-label AttacksChen Ma, Xiangyu Guo, Li Chen, Jun-Hai Yong 等NeurIPS 2021 · 被引用 24 次
- SoK: All You Need to Know About On-Device ML Model Extraction - The Gap Between Research and PracticeTushar Nayan, Qiming Guo, Mohammed Alduniawi, Marcus Botacin 等USENIX Security 2024 · 被引用 20 次
- Transferability of White-box Perturbations: Query-Efficient Adversarial Attacks against Commercial DNN ServicesMeng Shen, Changyue Li, Qi Li, Hao Lu 等USENIX Security 2024 · 被引用 8 次
- Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation ModelsVyas Raina, Rao Ma, Charles McGhee, Kate M. Knill 等EMNLP 2024 · 被引用 5 次
它引用的顶会 Paper8
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
- Stealing Hyperparameters in Machine LearningBinghui Wang, Neil Zhenqiang GongS&P 2018 · 被引用 504 次
- Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning AttacksAmbra Demontis, Marco Melis, Maura Pintor, Matthew Jagielski 等USENIX Security 2019 · 被引用 466 次
- Enhancing Adversarial Example Transferability With an Intermediate Level AttackQian Huang, Isay Katsman, Zeqi Gu, Horace He 等ICCV 2019 · 被引用 293 次
相关 Paper
- Defending Against Model Stealing Attacks With Adaptive MisinformationSanjay Kariyappa, Moinuddin K. QureshiCVPR 2020
- Query-efficient Meta Attack to Deep Neural NetworksJiawei Du, Hu Zhang, Joey Tianyi Zhou, Yi Yang 等ICLR 2020 · 被引用 87 次
- Delving into Data: Effectively Substitute Training for Black-box AttackWenxuan Wang, Bangjie Yin, Taiping Yao, Li Zhang 等CVPR 2021
- Training Meta-Surrogate Model for Transferable Adversarial AttackYunxiao Qin, Yuanhao Xiong, Jinfeng Yi, Cho-Jui HsiehAAAI 2023 · 被引用 31 次
- Protecting DNNs from Theft using an Ensemble of Diverse ModelsSanjay Kariyappa, Atul Prakash, Moinuddin K. QureshiICLR 2021 · 被引用 33 次
