On the Perils of Cascading Robust Classifiers
Ravi Mangal, Zifan Wang, Chi Zhang, Klas Leino, Corina S. Pasareanu, Matt Fredrikson
摘要
Ensembling certifiably robust neural networks is a promising approach for improving the certified robust accuracy of neural models. Black-box ensembles that assume only query-access to the constituent models (and their robustness certifiers) during prediction are particularly attractive due to their modular structure. Cascading ensembles are a popular instance of black-box ensembles that appear to improve certified robust accuracies in practice. However, we show that the robustness certifier used by a cascading ensemble is unsound. That is, when a cascading ensemble is certified as locally robust at an input x (with respect to ), there can be inputs x in the -ball centered at x, such that the cascade's prediction at x is different from x and thus the ensemble is not locally robust. Our theoretical findings are accompanied by empirical results that further demonstrate this unsoundness. We present cascade attack (CasA), an adversarial attack against cascading ensembles, and show that: (1) there exists an adversarial input for up to 88% of the samples where the ensemble claims to be certifiably robust and accurate; and (2) the accuracy of a cascading ensemble under our attack is as low as 11% when it claims to be certifiably robust and accurate on 97% of the test set. Our work reveals a critical pitfall of cascading certifiably robust models by showing that the seemingly beneficial strategy of cascading can actually hurt the robustness of the resulting ensemble. Our code is available at https://github.com/TristaChi/ensembleKW . * Equal Contribution 1 Percentage of inputs where the classifier is accurate and certified as locally robust.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Globally-Robust Neural NetworksKlas Leino, Zifan Wang, Matt FredriksonICML 2021 · 被引用 150 次
- DVERGE: Diversifying Vulnerabilities for Enhanced Robust Generation of EnsemblesHuanrui Yang, Jingyang Zhang, Hongliang Dong, Nathan Inkawhich 等NeurIPS 2020 · 被引用 144 次
- EMPIR: Ensembles of Mixed Precision Deep Networks for Increased Robustness Against Adversarial AttacksSanchari Sen, Balaraman Ravindran, Anand RaghunathanICLR 2020 · 被引用 69 次
- On the Certified Robustness for Ensemble Models and BeyondZhuolin Yang, Linyi Li, Xiaojun Xu, Bhavya Kailkhura 等ICLR 2022 · 被引用 57 次
- Fast Geometric Projections for Local Robustness CertificationAymeric Fromherz, Klas Leino, Matt Fredrikson, Bryan Parno 等ICLR 2021 · 被引用 34 次
相关 Paper
- Rethinking Model Ensemble in Transfer-based Adversarial AttacksHuanran Chen, Yichi Zhang, Yinpeng Dong, Xiao Yang 等ICLR 2024 · 被引用 112 次
- MORA: Improving Ensemble Robustness Evaluation with Model Reweighing AttackYunrui Yu, Xitong Gao, Cheng-Zhong XuNeurIPS 2022 · 被引用 14 次
- LEA2: A Lightweight Ensemble Adversarial Attack via Non-overlapping Vulnerable Frequency RegionsYaguan Qian, Shuke He, Chenyu Zhao, Jiaqiang Sha 等ICCV 2023 · 被引用 26 次
- Disrupting Deep Uncertainty Estimation Without Harming AccuracyIdo Galil, Ran El-YanivNeurIPS 2021 · 被引用 27 次
- Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable ConfidenceHanbin Hong, Xinyu Zhang, Binghui Wang, Zhongjie Ba 等CCS 2024 · 被引用 3 次
