Adversarial Robustness Guarantees for Random Deep Neural Networks
Giacomo De Palma, Bobak Toussi Kiani, Seth Lloyd
Abstract
The reliability of most deep learning algorithms is fundamentally challenged by the existence of adversarial examples, which are incorrectly classified inputs that are extremely close to a correctly classified input. We study adversarial examples for deep neural networks with random weights and biases and prove that the distance of any given input from the classification boundary scales at least as , where is the dimension of the input. We also extend our proof to cover all the norms. Our results constitute a fundamental advance in the study of adversarial examples, and encompass a wide variety of architectures, which include any combination of convolutional or fully connected layers with skipped connections and pooling. We validate our results with experiments on both random deep neural networks and deep neural networks trained on the MNIST and CIFAR10 datasets. Given the results of our experiments on MNIST and CIFAR10, we conjecture that the proof of our adversarial robustness guarantee can be extended to trained deep neural networks. This extension will open the way to a thorough theoretical study of neural network robustness by classifying the relation between network architecture and adversarial distance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28cfe60a-76b7-435e-8fb0-a4fab7408567Cited by top-tier papers4
- Fixed Neural Network Steganography: Train the images, not the networkVarsha Kishore, Xiangyu Chen, Yan Wang, Boyi Li et al.ICLR 2022 · 59 citations
- Planting Undetectable Backdoors in Machine Learning Models : [Extended Abstract]Shafi Goldwasser, Michael P. Kim, Vinod Vaikuntanathan, Or ZamirFOCS 2022 · 40 citations
- Evolution of Neural Tangent Kernels under Benign and Adversarial TrainingNoel Loo, Ramin M. Hasani, Alexander Amini, Daniela RusNeurIPS 2022 · 18 citations
- Adversarial Training from Mean Field PerspectiveSoichiro Kumano, Hiroshi Kera, Toshihiko YamasakiNeurIPS 2023 · 2 citations
Builds on5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 1,765 citations
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 167 citations
- Robustness of Bayesian Neural Networks to Gradient-Based AttacksGinevra Carbone, Matthew Wicker, Luca Laurenti, Andrea Patané et al.NeurIPS 2020 · 85 citations
Related papers
- Most ReLU Networks Suffer from Adversarial PerturbationsAmit Daniely, Hadas ShachamNeurIPS 2020 · 17 citations
- Feature compression is the root cause of adversarial fragility in neural networksJingchao Gao, Ziqing Lu, Raghu Mudumbai, Xiaodong Wu et al.ICLR 2026 · 3 citations
- Adversarial Examples in Multi-Layer Random ReLU NetworksPeter L. Bartlett, Sébastien Bubeck, Yeshwanth CherapanamjeriNeurIPS 2021 · 33 citations
- Generalization bounds for deep convolutional neural networksPhilip M. Long, Hanie SedghiICLR 2020 · 102 citations
- PAC-Bayesian Spectrally-Normalized Bounds for Adversarially Robust GeneralizationJiancong Xiao, Ruoyu Sun, Zhi-Quan LuoNeurIPS 2023 · 14 citations
