Almost Tight L0-norm Certified Robustness of Top-k Predictions against Adversarial Perturbations
Jinyuan Jia, Binghui Wang, Xiaoyu Cao, Hongbin Liu, Neil Zhenqiang Gong
Abstract
Top- predictions are used in many real-world applications such as machine learning as a service, recommender systems, and web searches. -norm adversarial perturbation characterizes an attack that arbitrarily modifies some features of an input such that a classifier makes an incorrect prediction for the perturbed input. -norm adversarial perturbation is easy to interpret and can be implemented in the physical world. Therefore, certifying robustness of top- predictions against -norm adversarial perturbation is important. However, existing studies either focused on certifying -norm robustness of top- predictions or -norm robustness of top- predictions. In this work, we aim to bridge the gap. Our approach is based on randomized smoothing, which builds a provably robust classifier from an arbitrary classifier via randomizing an input. Our major theoretical contribution is an almost tight -norm certified robustness guarantee for top- predictions. We empirically evaluate our method on CIFAR10 and ImageNet. For instance, our method can build a classifier that achieves a certified top-3 accuracy of 69.2% on ImageNet when an attacker can arbitrarily perturb 5 pixels of a testing image.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36e1eaba-91c2-4e3c-b657-bf0b670dfa03Cited by top-tier papers9
- RS-Del: Edit Distance Robustness Certificates for Sequence Classifiers via Randomized DeletionZhuoqun Huang, Neil G. Marchant, Keane Lucas, Lujo Bauer et al.NeurIPS 2023 · 24 citations
- MultiGuard: Provably Robust Multi-label Classification against Adversarial ExamplesJinyuan Jia, Wenjie Qu, Neil Zhenqiang GongNeurIPS 2022 · 22 citations
- Node-aware Bi-smoothing: Certified Robustness against Graph Injection AttacksYuni Lai, Yulin Zhu, Bailin Pan, Kai ZhouS&P 2024 · 11 citations
- MMCert: Provable Defense Against Adversarial Attacks to Multi-Modal ModelsYanting Wang, Hongye Fu, Wei Zou, Jinyuan JiaCVPR 2024 · 4 citations
- Collective Certified Robustness against Graph Injection AttacksYuni Lai, Bailin Pan, Kaihuang Chen, Yancheng Yuan et al.ICML 2024 · 4 citations
Builds on13
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- AI2: Safety and Robustness Certification of Neural Networks with Abstract InterpretationTimon Gehr, Matthew Mirman, Dana Drachsler-Cohen, Petar Tsankov et al.S&P 2018 · 987 citations
- Formal Security Analysis of Neural Networks using Symbolic IntervalsShiqi Wang, Kexin Pei, Justin Whitehouse, Junfeng Yang et al.USENIX Security 2018 · 523 citations
- Randomized Smoothing of All Shapes and SizesGreg Yang, Tony Duan, J. Edward Hu, Hadi Salman et al.ICML 2020 · 237 citations
- MACER: Attack-free and Scalable Robust Training via Maximizing Certified RadiusRuntian Zhai, Chen Dan, Di He, Huan Zhang et al.ICLR 2020 · 195 citations
Related papers
- Certified Robustness for Top-k Predictions against Adversarial Perturbations via Randomized SmoothingJinyuan Jia, Xiaoyu Cao, Binghui Wang, Neil Zhenqiang GongICLR 2020 · 107 citations
- Robustness Certificates for Sparse Adversarial Attacks by Randomized AblationAlexander Levine, Soheil FeiziAAAI 2020 · 114 citations
- Higher-Order Certification For Randomized SmoothingJeet Mohapatra, Ching-Yun Ko, Tsui-Wei Weng, Pin-Yu Chen et al.NeurIPS 2020 · 51 citations
- Improving l1-Certified Robustness via Randomized Smoothing by Leveraging Box ConstraintsVáclav Vorácek, Matthias HeinICML 2023 · 11 citations
- Certified Robustness to Label-Flipping Attacks via Randomized SmoothingElan Rosenfeld, Ezra Winston, Pradeep Ravikumar, J. Zico KolterICML 2020 · 182 citations
