MORA: Improving Ensemble Robustness Evaluation with Model Reweighing Attack
Yunrui Yu, Xitong Gao, Cheng-Zhong Xu
Abstract
Adversarial attacks can deceive neural networks by adding tiny perturbations to their input data. Ensemble defenses, which are trained to minimize attack transferability among sub-models, offer a promising research direction to improve robustness against such attacks while maintaining a high accuracy on natural inputs. We discover, however, that recent state-of-the-art (SOTA) adversarial attack strategies cannot reliably evaluate ensemble defenses, sizeably overestimating their robustness. This paper identifies the two factors that contribute to this behavior. First, these defenses form ensembles that are notably difficult for existing gradient-based method to attack, due to gradient obfuscation. Second, ensemble defenses diversify sub-model gradients, presenting a challenge to defeat all sub-models simultaneously, simply summing their contributions may counteract the overall attack objective; yet, we observe that ensemble may still be fooled despite most sub-models being correct. We therefore introduce MORA, a model-reweighing attack to steer adversarial example synthesis by reweighing the importance of sub-model gradients. MORA finds that recent ensemble defenses all exhibit varying degrees of overestimated robustness. Comparing it against recent SOTA white-box attacks, it can converge orders of magnitude faster while achieving higher attack success rates across all ensemble models examined with three different ensemble modes (i.e., ensembling by either softmax, voting or logits). In particular, most ensemble defenses exhibit near or exactly 0% robustness against MORA with perturbation within 0.02 on CIFAR-10, and 0.01 on CIFAR-100. We make MORA open source with reproducible results and pre-trained models; and provide a leaderboard of ensemble defenses under various attack strategies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21fbcce1-5198-405a-8ae8-8c2fd1d7a585Cited by top-tier papers2
- AdvDiffuser: Natural Adversarial Example Synthesis with Diffusion ModelsXinquan Chen, Xitong Gao, Juanjuan Zhao, Kejiang Ye et al.ICCV 2023 · 94 citations
- Stop Diverse OOD Attacks: Knowledge Ensemble for Reliable DefenseZhenbo Shi, Xiaoman Liu, Yuxuan Zhang, Shuchang Wang et al.AAAI 2025
Builds on16
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Data Augmentation Can Improve RobustnessSylvestre-Alvise Rebuffi, Sven Gowal, Dan Andrei Calian, Florian Stimberg et al.NeurIPS 2021 · 427 citations
- Improving Robustness using Generated DataSven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles, Florian Stimberg et al.NeurIPS 2021 · 384 citations
Related papers
- LEA2: A Lightweight Ensemble Adversarial Attack via Non-overlapping Vulnerable Frequency RegionsYaguan Qian, Shuke He, Chenyu Zhao, Jiaqiang Sha et al.ICCV 2023 · 26 citations
- Strong Transferable Adversarial Attacks via Ensembled Asymptotically Normal Distribution LearningZhengwei Fang, Rui Wang, Tao Huang, Liping JingCVPR 2024
- Admix: Enhancing the Transferability of Adversarial AttacksXiaosen Wang, Xuanran He, Jingdong Wang, Kun HeICCV 2021 · 282 citations
- On the Robustness of Vision Transformers to Adversarial ExamplesKaleel Mahmood, Rigel Mahmood, Marten van DijkICCV 2021 · 261 citations
- Ensemble Diversity Facilitates Adversarial TransferabilityBowen Tang, Zheng Wang, Yi Bin, Qi Dou et al.CVPR 2024 · 22 citations
