DVERGE: Diversifying Vulnerabilities for Enhanced Robust Generation of Ensembles
Huanrui Yang, Jingyang Zhang, Hongliang Dong, Nathan Inkawhich, Andrew Gardner, Andrew Touchet, Wesley Wilkes, Heath Berry, Hai Li
Abstract
Recent research finds CNN models for image classification demonstrate overlapped adversarial vulnerabilities: adversarial attacks can mislead CNN models with small perturbations, which can effectively transfer between different models trained on the same dataset. Adversarial training, as a general robustness improvement technique, eliminates the vulnerability in a single model by forcing it to learn robust features. The process is hard, often requires models with large capacity, and suffers from significant loss on clean data accuracy. Alternatively, ensemble methods are proposed to induce sub-models with diverse outputs against a transfer adversarial example, making the ensemble robust against transfer attacks even if each sub-model is individually non-robust. Only small clean accuracy drop is observed in the process. However, previous ensemble training methods are not efficacious in inducing such diversity and thus ineffective on reaching robust ensemble. We propose DVERGE, which isolates the adversarial vulnerability in each sub-model by distilling non-robust features, and diversifies the adversarial vulnerability to induce diverse outputs against a transfer attack. The novel diversity metric and training procedure enables DVERGE to achieve higher robustness against transfer attacks comparing to previous ensemble methods, and enables the improved robustness when more sub-models are added to the ensemble. The code of this work is available at https://github.com/zjysteven/DVERGE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dcd118a7-b4a5-4e97-82c8-5cf12bba4b3aCited by top-tier papers35
- Cross-Entropy Loss Functions: Theoretical Analysis and ApplicationsAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2023 · 790 citations
- An Adaptive Model Ensemble Adversarial Attack for Boosting Adversarial TransferabilityBin Chen, Jia-Li Yin, Shukai Chen, Bohao Chen et al.ICCV 2023 · 96 citations
- 3D Common Corruptions and Data AugmentationOguzhan Fatih Kar, Teresa Yeo, Andrei Atanov, Amir ZamirCVPR 2022 · 80 citations
- Random Noise Defense Against Query-Based Black-Box AttacksZeyu Qin, Yanbo Fan, Hongyuan Zha, Baoyuan WuNeurIPS 2021 · 78 citations
- On the Certified Robustness for Ensemble Models and BeyondZhuolin Yang, Linyi Li, Xiaojun Xu, Bhavya Kailkhura et al.ICLR 2022 · 57 citations
Builds on9
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 569 citations
- Skip Connections Matter: On the Transferability of Adversarial Examples Generated with ResNetsDongxian Wu, Yisen Wang, Shu-Tao Xia, James Bailey et al.ICLR 2020 · 357 citations
Related papers
- To Tackle Adversarial Transferability: A Novel Ensemble Training Method with Fourier TransformationWanlin Zhang, Weichen Lin, Ruomin Huang, Shihong Song et al.ICLR 2025
- Adversarial Defence by Diversified Simultaneous Training of Deep EnsemblesBo Huang, Zhiwei Ke, Yi Wang, Wei Wang et al.AAAI 2021 · 20 citations
- Improving Ensemble Robustness by Collaboratively Promoting and Demoting Adversarial RobustnessTuan-Anh Bui, Trung Le, He Zhao, Paul Montague et al.AAAI 2021 · 13 citations
- Rethinking Model Ensemble in Transfer-based Adversarial AttacksHuanran Chen, Yichi Zhang, Yinpeng Dong, Xiao Yang et al.ICLR 2024 · 112 citations
- Transferable Perturbations of Deep Feature DistributionsNathan Inkawhich, Kevin J. Liang, Lawrence Carin, Yiran ChenICLR 2020 · 100 citations
