Dual Manifold Adversarial Robustness: Defense against Lp and non-Lp Adversarial Attacks
Wei-An Lin, Chun Pong Lau, Alexander Levine, Rama Chellappa, Soheil Feizi
Abstract
Adversarial training is a popular defense strategy against attack threat models with bounded L p norms. However, it often degrades the model performance on normal images and more importantly, the defense does not generalize well to novel attacks. Given the success of deep generative models such as GANs and VAEs in characterizing (approximately) the underlying manifold of images, we investigate whether or not the aforementioned deficiencies of adversarial training can be remedied by exploiting the underlying manifold information. To partially answer this question, we consider the scenario when the manifold information of the underlying data is available. We use a subset of ImageNet natural images where an approximate underlying manifold is learned using StyleGAN. We also construct an "On-Manifold ImageNet" (OM-ImageNet) dataset by projecting the ImageNet samples onto the learned manifold. For this dataset, the underlying manifold information is exact. Using OM-ImageNet, we first show that adversarial training in the latent space of images (i.e. on-manifold adversarial training) improves both standard accuracy and robustness to on-manifold attacks. However, since no out-of-manifold perturbations are realized, the defense can be broken by L p adversarial attacks. We further propose Dual Manifold Adversarial Training (DMAT) where adversarial perturbations in both latent and image spaces are used in robustifying the model. Our DMAT improves performance on normal images, and achieves comparable robustness to the standard adversarial training against L p attacks. In addition, we observe that models defended by DMAT achieve improved robustness against novel attacks which manipulate images by global color shifts or various types of image filtering. Interestingly, similar improvements are also achieved when the defended models are tested on (out-of-manifold) natural images. These results demonstrate the potential benefits of using manifold information (exactly or approximately) in enhancing robustness of deep learning models against various types of novel adversarial attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- DISCO: Adversarial Defense with Local Implicit FunctionsChih-Hui Ho, Nuno VasconcelosNeurIPS 2022 · 65 citations
- Why Does Little Robustness Help? A Further Step Towards Understanding Adversarial TransferabilityYechao Zhang, Shengshan Hu, Leo Yu Zhang, Junyu Shi et al.S&P 2024 · 36 citations
- Explicit Tradeoffs between Adversarial and Natural Distributional RobustnessMazda Moayeri, Kiarash Banihashem, Soheil FeiziNeurIPS 2022 · 28 citations
- MixACM: Mixup-Based Robustness Transfer via Distillation of Activated Channel MapsMuhammad Awais, Fengwei Zhou, Chuanlong Xie, Jiawei Li et al.NeurIPS 2021 · 22 citations
- A Theory of Transfer-Based Black-Box Attacks: Explanation and ImplicationsYanbo Chen, Weiwei LiuNeurIPS 2023 · 22 citations
Builds on9
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
Related papers
- Deep Manifold Attack on Point Clouds via Parameter Plane StretchingKeke Tang, Jianpeng Wu, Weilong Peng, Yawen Shi et al.AAAI 2023 · 25 citations
- Achieving Robustness in the Wild via Adversarial Mixing With Disentangled RepresentationsSven Gowal, Chongli Qin, Po-Sen Huang, A. Taylan Cemgil et al.CVPR 2020
- Robustness and Generalization via Generative Adversarial TrainingOmid Poursaeed, Tianxing Jiang, Harry Yang, Serge J. Belongie et al.ICCV 2021 · 35 citations
- Boosting Adversarial Training with Hypersphere EmbeddingTianyu Pang, Xiao Yang, Yinpeng Dong, Taufik Xu et al.NeurIPS 2020 · 170 citations
- Exploring Adversarial Fake Images on Face ManifoldDongze Li, Wei Wang, Hongxing Fan, Jing DongCVPR 2021
