Enhancing the Adversarial Robustness via Manifold Projection
Zhiting Li, Shibai Yin, Tai-Xiang Jiang, Yexun Hu, Jia-Mian Wu, Guowei Yang, Guisong Liu
Abstract
Deep learning has been widely applied to various aspects of computer vision, but the emergence of adversarial attacks raises concerns about its reliability. Adversarial training (AT) is one of the most effective defense methods, which incorporates adversarial examples into the training data. However, AT is typically employed in a discriminative learning manner, i.e., learning the mapping (conditional probability) from samples to labels, it essentially reinforces this mapping without considering the underlying data distribution. It is notable that adversarial examples often deviate from the distribution of normal (clean) samples. Therefore, building upon existing adversarial defense schemes, we propose to further exploit the distribution of normal samples, partly from the generative learning perspective, resulting in a novel robustness enhancement paradigm. We train a simple autoencoder (AE) autoregressively on normal samples to learn their prior distribution, effectively serving as an image manifold. This AE is then used as a manifold projection operator to incorporate the distribution information of normal samples. Specifically, we organically integrate the pretrained AE into the training process of both AT and adversarial distillation (AD), a method aiming at improving the robustness of small models with low capacity. Since the AE captures the distribution of normal samples, it can adaptively pull adversarial examples closer to the normal sample manifold, weakening the attack strength of adversarial samples and easing the learning of mappings from adversarial samples to correct labels. From the Pearson correlation coefficient (PCC) between the statistics on normal and adversarial examples, it’s validated that the AE indeed pulls adversarial samples closer to normal samples. Extensive experiments illustrate that our proposed adversarial defense paradigm significantly improves the robustness compared with previous state-of-the-art AT and AD methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a12e979b-430f-4675-bfe5-727fb60706d8Cited by top-tier papers1
Ask how each one uses itBuilds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
Related papers
- Towards Robustness of Deep Neural Networks via RegularizationYao Li, Martin Renqiang Min, Thomas C. M. Lee, Wenchao Yu et al.ICCV 2021 · 8 citations
- Adversarial Distributional Training for Robust Deep LearningYinpeng Dong, Zhijie Deng, Tianyu Pang, Jun Zhu et al.NeurIPS 2020 · 154 citations
- Improving Adversarial Robustness via Guided Complement EntropyHao-Yun Chen, Jhao-Hong Liang, Shih-Chieh Chang, Jia-Yu Pan et al.ICCV 2019 · 51 citations
- Efficient Adversarial Training With Transferable Adversarial ExamplesHaizhong Zheng, Ziqi Zhang, Juncheng Gu, Honglak Lee et al.CVPR 2020
- Adversarial Purification with the Manifold HypothesisZhaoyuan Yang, Zhiwei Xu, Jing Zhang, Richard I. Hartley et al.AAAI 2024 · 12 citations
