Adversarial Neural Pruning with Latent Vulnerability Suppression
Divyam Madaan, Jinwoo Shin, Sung Ju Hwang
Abstract
Despite the remarkable performance of deep neural networks on various computer vision tasks, they are known to be susceptible to adversarial perturbations, which makes it challenging to deploy them in real-world safety-critical applications. In this paper, we conjecture that the leading cause of adversarial vulnerability is the distortion in the latent feature space, and provide methods to suppress them effectively. Explicitly, we define vulnerability for each latent feature and then propose a new loss for adversarial learning, Vulnerability Suppression (VS) loss, that aims to minimize the feature-level vulnerability during training. We further propose a Bayesian framework to prune features with high vulnerability to reduce both vulnerability and loss on adversarial samples. We validate our Adversarial Neural Pruning with Vulnerability Suppression (ANP-VS) method on multiple benchmark datasets, on which it not only obtains state-of-the-art adversarial robustness but also improves the performance on clean examples, using only a fraction of the parameters used by the full network. Further qualitative analysis suggests that the improvements come from the suppression of feature-level vulnerability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Adversarial Self-Supervised Contrastive LearningMinseon Kim, Jihoon Tack, Sung Ju HwangNeurIPS 2020 · 294 citations
- Improving Adversarial Robustness via Channel-wise Activation SuppressingYang Bai, Yuyuan Zeng, Yong Jiang, Shu-Tao Xia et al.ICLR 2021 · 59 citations
- Drawing Robust Scratch Tickets: Subnetworks with Inborn Robustness Are Found within Randomly Initialized NetworksYonggan Fu, Qixuan Yu, Yang Zhang, Shang Wu et al.NeurIPS 2021 · 36 citations
- Training Adversarially Robust Sparse Networks via Bayesian Connectivity SamplingOzan Özdenizci, Robert LegensteinICML 2021 · 31 citations
- Learning to Generate Noise for Multi-Attack RobustnessDivyam Madaan, Jinwoo Shin, Sung Ju HwangICML 2021 · 31 citations
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- Adversarial Robustness vs. Model Compression, or Both?Shaokai Ye, Xue Lin, Kaidi Xu, Sijia Liu et al.ICCV 2019 · 180 citations
Related papers
- Masking Adversarial Damage: Finding Adversarial Saliency for Robust and Sparse NetworkByung-Kwan Lee, Junho Kim, Yong Man RoCVPR 2022 · 8 citations
- Learn2Perturb: An End-to-End Feature Perturbation Learning to Improve Adversarial RobustnessAhmadreza Jeddi, Mohammad Javad Shafiee, Michelle Karg, Christian Scharfenberger et al.CVPR 2020
- Improving Adversarial Robustness via Probabilistically Compact Loss with Logit ConstraintsXin Li, Xiangrui Li, Deng Pan, Dongxiao ZhuAAAI 2021 · 17 citations
- Adversarial Robustness Through the Lens of CausalityYonggang Zhang, Mingming Gong, Tongliang Liu, Gang Niu et al.ICLR 2022 · 65 citations
- Nasty Adversarial Training: A Probability Sparsity Perspective for Robustness EnhancementYuhang Zhou, Zhongyun Hua, Zhaoquan Gu, Keke Tang et al.ICLR 2026
