Adversarial Example Games
Avishek Joey Bose, Gauthier Gidel, Hugo Berard, Andre Cianflone, Pascal Vincent, Simon Lacoste-Julien, William L. Hamilton
Abstract
The existence of adversarial examples capable of fooling trained neural network classifiers calls for a much better understanding of possible attacks to guide the development of safeguards against them. This includes attack methods in the challenging non-interactive blackbox setting, where adversarial attacks are generated without any access, including queries, to the target model. Prior attacks in this setting have relied mainly on algorithmic innovations derived from empirical observations (e.g., that momentum helps), lacking principled transferability guarantees. In this work, we provide a theoretical foundation for crafting transferable adversarial examples to entire hypothesis classes. We introduce Adversarial Example Games (AEG), a framework that models the crafting of adversarial examples as a min-max game between a generator of attacks and a classifier. AEG provides a new way to design adversarial examples by adversarially training a generator and a classifier from a given hypothesis class (e.g., architecture). We prove that this game has an equilibrium, and that the optimal generator is able to craft adversarial examples that can attack any classifier from the corresponding hypothesis class. We demonstrate the efficacy of AEG on the MNIST and CIFAR-10 datasets, outperforming prior state-of-the-art approaches with an average relative improvement of 29.9% and 47.2% against undefended and robust models (Table 2 & 3) respectively. * Equal Contribution, order chosen via randomization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 221e9f18-d53f-4e24-af7c-d477cfe9f9caCited by top-tier papers12
- Backdoor Attack with Imperceptible Input and Latent ModificationKhoa D. Doan, Yingjie Lao, Ping LiNeurIPS 2021 · 179 citations
- Adversarial Attack Generation Empowered by Min-Max OptimizationJingkang Wang, Tianyun Zhang, Sijia Liu, Pin-Yu Chen et al.NeurIPS 2021 · 49 citations
- Clipped Stochastic Methods for Variational Inequalities with Heavy-Tailed NoiseEduard Gorbunov, Marina Danilova, David Dobre, Pavel E. Dvurechenskii et al.NeurIPS 2022 · 36 citations
- Training Meta-Surrogate Model for Transferable Adversarial AttackYunxiao Qin, Yuanhao Xiong, Jinfeng Yi, Cho-Jui HsiehAAAI 2023 · 31 citations
- Balance, Imbalance, and Rebalance: Understanding Robust Overfitting from a Minimax Game PerspectiveYifei Wang, Liangchen Li, Jiansheng Yang, Zhouchen Lin et al.NeurIPS 2023 · 26 citations
Builds on9
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Minimally distorted Adversarial Examples with a Fast Adaptive Boundary AttackFrancesco Croce, Matthias HeinICML 2020 · 597 citations
- Skip Connections Matter: On the Transferability of Adversarial Examples Generated with ResNetsDongxian Wu, Yisen Wang, Shu-Tao Xia, James Bailey et al.ICLR 2020 · 357 citations
- MMA Training: Direct Input Space Margin Maximization through Adversarial TrainingGavin Weiguang Ding, Yash Sharma, Kry Yik Chau Lui, Ruitong HuangICLR 2020 · 308 citations
Related papers
- Admix: Enhancing the Transferability of Adversarial AttacksXiaosen Wang, Xuanran He, Jingdong Wang, Kun HeICCV 2021 · 282 citations
- Meta Gradient Adversarial AttackZheng Yuan, Jie Zhang, Yunpei Jia, Chuanqi Tan et al.ICCV 2021 · 95 citations
- Natural Color Fool: Towards Boosting Black-box Unrestricted AttacksShengming Yuan, Qilong Zhang, Lianli Gao, Yaya Cheng et al.NeurIPS 2022 · 86 citations
- A Little Robustness Goes a Long Way: Leveraging Robust Features for Targeted Transfer AttacksJacob M. Springer, Melanie Mitchell, Garrett T. KenyonNeurIPS 2021 · 54 citations
- MGAAttack: Toward More Query-efficient Black-box Attack by Microbial Genetic AlgorithmLina Wang, Kang Yang, Wenqi Wang, Run Wang et al.ACM MM 2020 · 9 citations
