On Breaking Deep Generative Model-based Defenses and Beyond
Yanzhi Chen, Renjie Xie, Zhanxing Zhu
Abstract
Deep neural networks have been proven to be vulnerable to the so-called adversarial attacks. Recently there have been efforts to defend such attacks with deep generative models. These defenses often predict by inverting the deep generative models rather than simple feedforward propagation. Such defenses are difficult to attack due to the obfuscated gradients caused by inversion. In this work, we propose a new white-box attack to break these defenses. The idea is to view the inversion phase as a dynamical system, through which we extract the gradient w.r.t the image by backtracking its trajectory. An amortized strategy is also developed to accelerate the attack. Experiments show that our attack better breaks stateof-the-art defenses (e.g DefenseGAN, ABS) than other attacks (e.g BPDA). Additionally, our empirical results provide insights for understanding the weaknesses of deep generative model defenses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- Eliminating Adversarial Noise via Information Discard and Robust Representation RestorationDawei Zhou, Yukun Chen, Nannan Wang, Decheng Liu et al.ICML 2023 · 10 citations
- Post-breach Recovery: Protection against White-box Adversarial Examples for Leaked DNN ModelsShawn Shan, Wenxin Ding, Emily Wenger, Haitao Zheng et al.CCS 2022 · 9 citations
- AdvFlow: Inconspicuous Black-box Adversarial Attacks using Normalizing FlowsHadi Mohaghegh Dolatabadi, Sarah M. Erfani, Christopher LeckieNeurIPS 2020 · 75 citations
- Ensemble Generative Cleaning With Feedback Loops for Defending Adversarial AttacksJianhe Yuan, Zhihai HeCVPR 2020
- Understanding the Robustness of Randomized Feature Defense Against Query-Based Adversarial AttacksNguyen Hung-Quang, Yingjie Lao, Tung Pham, Kok-Seng Wong et al.ICLR 2024 · 3 citations
