On the Perils of Cascading Robust Classifiers
Ravi Mangal, Zifan Wang, Chi Zhang, Klas Leino, Corina S. Pasareanu, Matt Fredrikson
Abstract
Ensembling certifiably robust neural networks is a promising approach for improving the certified robust accuracy of neural models. Black-box ensembles that assume only query-access to the constituent models (and their robustness certifiers) during prediction are particularly attractive due to their modular structure. Cascading ensembles are a popular instance of black-box ensembles that appear to improve certified robust accuracies in practice. However, we show that the robustness certifier used by a cascading ensemble is unsound. That is, when a cascading ensemble is certified as locally robust at an input x (with respect to ), there can be inputs x in the -ball centered at x, such that the cascade's prediction at x is different from x and thus the ensemble is not locally robust. Our theoretical findings are accompanied by empirical results that further demonstrate this unsoundness. We present cascade attack (CasA), an adversarial attack against cascading ensembles, and show that: (1) there exists an adversarial input for up to 88% of the samples where the ensemble claims to be certifiably robust and accurate; and (2) the accuracy of a cascading ensemble under our attack is as low as 11% when it claims to be certifiably robust and accurate on 97% of the test set. Our work reveals a critical pitfall of cascading certifiably robust models by showing that the seemingly beneficial strategy of cascading can actually hurt the robustness of the resulting ensemble. Our code is available at https://github.com/TristaChi/ensembleKW . * Equal Contribution 1 Percentage of inputs where the classifier is accurate and certified as locally robust.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c212ef3-35c6-4d37-8e6a-3180765dcfe4Builds on8
- Globally-Robust Neural NetworksKlas Leino, Zifan Wang, Matt FredriksonICML 2021 · 150 citations
- DVERGE: Diversifying Vulnerabilities for Enhanced Robust Generation of EnsemblesHuanrui Yang, Jingyang Zhang, Hongliang Dong, Nathan Inkawhich et al.NeurIPS 2020 · 144 citations
- EMPIR: Ensembles of Mixed Precision Deep Networks for Increased Robustness Against Adversarial AttacksSanchari Sen, Balaraman Ravindran, Anand RaghunathanICLR 2020 · 69 citations
- On the Certified Robustness for Ensemble Models and BeyondZhuolin Yang, Linyi Li, Xiaojun Xu, Bhavya Kailkhura et al.ICLR 2022 · 57 citations
- Fast Geometric Projections for Local Robustness CertificationAymeric Fromherz, Klas Leino, Matt Fredrikson, Bryan Parno et al.ICLR 2021 · 34 citations
Related papers
- Rethinking Model Ensemble in Transfer-based Adversarial AttacksHuanran Chen, Yichi Zhang, Yinpeng Dong, Xiao Yang et al.ICLR 2024 · 112 citations
- MORA: Improving Ensemble Robustness Evaluation with Model Reweighing AttackYunrui Yu, Xitong Gao, Cheng-Zhong XuNeurIPS 2022 · 14 citations
- LEA2: A Lightweight Ensemble Adversarial Attack via Non-overlapping Vulnerable Frequency RegionsYaguan Qian, Shuke He, Chenyu Zhao, Jiaqiang Sha et al.ICCV 2023 · 26 citations
- Disrupting Deep Uncertainty Estimation Without Harming AccuracyIdo Galil, Ran El-YanivNeurIPS 2021 · 27 citations
- Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable ConfidenceHanbin Hong, Xinyu Zhang, Binghui Wang, Zhongjie Ba et al.CCS 2024 · 3 citations
