Clustering Effect of Adversarial Robust Models
Yang Bai, Xin Yan, Yong Jiang, Shu-Tao Xia, Yisen Wang
Abstract
Adversarial robustness has received increasing attention along with the study of adversarial examples. So far, existing works show that robust models not only obtain robustness against various adversarial attacks but also boost the performance in some downstream tasks. However, the underlying mechanism of adversarial robustness is still not clear. In this paper, we interpret adversarial robustness from the perspective of linear components, and find that there exist some statistical properties for comprehensively robust models. Specifically, robust models show obvious hierarchical clustering effect on their linearized sub-networks, when removing or replacing all non-linear components (e.g., batch normalization, maximum pooling, or activation layers). Based on these observations, we propose a novel understanding of adversarial robustness and apply it on more tasks including domain adaption and robustness boosting. Experimental evaluations demonstrate the rationality and superiority of our proposed clustering strategy. Our code is available at https://github.com/bymavis/Adv_Weight_NeurIPS2021.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8dcc487f-64c1-4f9c-9e44-1b910af74219Cited by top-tier papers8
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial TrainingLue Tao, Lei Feng, Jinfeng Yi, Sheng-Jun Huang et al.NeurIPS 2021 · 90 citations
- What Can the Neural Tangent Kernel Tell Us About Adversarial Robustness?Nikolaos Tsilivis, Julia KempeNeurIPS 2022 · 28 citations
- A Unified Contrastive Energy-based Model for Understanding the Generative Ability of Adversarial TrainingYifei Wang, Yisen Wang, Jiansheng Yang, Zhouchen LinICLR 2022 · 19 citations
- On the Functional Similarity of Robust and Non-Robust Neural RepresentationsAndrás Balogh, Márk JelasityICML 2023 · 4 citations
- Enhancing Robustness in Incremental Learning with Adversarial TrainingSeungju Cho, Hongsin Lee, Changick KimAAAI 2025 · 4 citations
Builds on10
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Do Adversarially Robust ImageNet Models Transfer Better?Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor et al.NeurIPS 2020 · 506 citations
Related papers
- Intriguing Properties of Adversarial Training at ScaleCihang Xie, Alan L. YuilleICLR 2020 · 66 citations
- Towards a Unified Game-Theoretic View of Adversarial Perturbations and RobustnessJie Ren, Die Zhang, Yisen Wang, Lu Chen et al.NeurIPS 2021 · 27 citations
- Explaining Adversarial Robustness of Neural Networks from Clustering Effect PerspectiveYulin Jin, Xiaoyu Zhang, Jian Lou, Xu Ma et al.ICCV 2023 · 3 citations
- To Tackle Adversarial Transferability: A Novel Ensemble Training Method with Fourier TransformationWanlin Zhang, Weichen Lin, Ruomin Huang, Shihong Song et al.ICLR 2025
- PAC-Bayesian Spectrally-Normalized Bounds for Adversarially Robust GeneralizationJiancong Xiao, Ruoyu Sun, Zhi-Quan LuoNeurIPS 2023 · 14 citations
