Breaking Certified Defenses: Semantic Adversarial Examples with Spoofed robustness Certificates
Amin Ghiasi, Ali Shafahi, Tom Goldstein
Abstract
Defenses against adversarial attacks can be classified into certified and non-certified. Certifiable defenses make networks robust within a certain -bounded radius, so that it is impossible for the adversary to make adversarial examples in the certificate bound. We present an attack that maintains the imperceptibility property of adversarial examples while being outside of the certified radius. Furthermore, the proposed "Shadow Attack" can fool certifiably robust networks by producing an imperceptible adversarial example that gets misclassified and produces a strong ``spoofed'' certificate.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70c3624d-3387-4500-9060-ef46d2d4cb15Cited by top-tier papers9
- MetaPoison: Practical General-purpose Clean-label Data PoisoningW. Ronny Huang, Jonas Geiping, Liam Fowl, Gavin Taylor et al.NeurIPS 2020 · 242 citations
- Fixed Neural Network Steganography: Train the images, not the networkVarsha Kishore, Xiangyu Chen, Yan Wang, Boyi Li et al.ICLR 2022 · 59 citations
- Prompt Certified Machine Unlearning with Randomized Gradient Smoothing and QuantizationZijie Zhang, Yang Zhou, Xin Zhao, Tianshi Che et al.NeurIPS 2022 · 56 citations
- TSS: Transformation-Specific Smoothing for Robustness CertificationLinyi Li, Maurice Weber, Xiaojun Xu, Luka Rimanic et al.CCS 2021 · 21 citations
- Why adversarial training can hurt robust accuracyJacob Clarysse, Julia Hörrmann, Fanny YangICLR 2023 · 6 citations
Builds on3
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal et al.ICLR 2020 · 384 citations
- Universal Adversarial TrainingAli Shafahi, Mahyar Najibi, Zheng Xu, John P. Dickerson et al.AAAI 2020 · 210 citations
Related papers
- Certified but Fooled! Breaking Certified Defenses with Ghost CertificatesViet Quoc Vo, Tashreque Mohammed Haq, Paul Montague, Tamas Abraham et al.AAAI 2026
- Deterministic Certification of Graph Neural Networks against Graph Poisoning Attacks with Arbitrary PerturbationsJiate Li, Meng Pang, Yun Dong, Binghui WangCVPR 2025
- Computational Asymmetries in Robust ClassificationSamuele Marro, Michele LombardiICML 2023 · 2 citations
- (De)Randomized Smoothing for Certifiable Defense against Patch AttacksAlexander Levine, Soheil FeiziNeurIPS 2020 · 188 citations
- Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial PerturbationsFlorian Tramèr, Jens Behrmann, Nicholas Carlini, Nicolas Papernot et al.ICML 2020 · 103 citations
