Robustness in deep learning: The good (width), the bad (depth), and the ugly (initialization)
Zhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Volkan Cevher
Abstract
We study the average robustness notion in deep neural networks in (selected) wide and narrow, deep and shallow, as well as lazy and non-lazy training settings. We prove that in the under-parameterized setting, width has a negative effect while it improves robustness in the over-parameterized setting. The effect of depth closely depends on the initialization and the training mode. In particular, when initialized with LeCun initialization, depth helps robustness with the lazy training regime. In contrast, when initialized with Neural Tangent Kernel (NTK) and He-initialization, depth hurts the robustness. Moreover, under the non-lazy training regime, we demonstrate how the width of a two-layer ReLU network benefits robustness. Our theoretical developments improve the results by [Huang et al. NeurIPS21; Wu et al. NeurIPS21] and are consistent with [Bubeck and Sellke NeurIPS21; Bubeck et al. COLT21].
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 826fae29-ce3a-4a8c-8382-a1c8f38ef559Cited by top-tier papers10
- Sample Complexity Bounds for Score-Matching: Causal Discovery and Generative ModelingZhenyu Zhu, Francesco Locatello, Volkan CevherNeurIPS 2023 · 23 citations
- Initialization Matters: Privacy-Utility Analysis of Overparameterized Neural NetworksJiayuan Ye, Zhenyu Zhu, Fanghui Liu, Reza Shokri et al.NeurIPS 2023 · 19 citations
- GlucoBench: Curated List of Continuous Glucose Monitoring Datasets with Prediction BenchmarksRenat Sergazinov, Elizabeth Chun, Valeriya Rogovchenko, Nathaniel J. Fernandes et al.ICLR 2024 · 13 citations
- Beyond the Universal Law of Robustness: Sharper Laws for Random Features and Neural Tangent KernelsSimone Bombari, Shayan Kiyani, Marco MondelliICML 2023 · 13 citations
- Robust NAS under adversarial training: benchmark, theory, and beyondYongtao Wu, Fanghui Liu, Carl-Johann Simon-Gabriel, Grigorios Chrysos et al.ICLR 2024 · 10 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Proving the Lottery Ticket Hypothesis: Pruning is All You NeedEran Malach, Gilad Yehudai, Shai Shalev-Shwartz, Ohad ShamirICML 2020 · 327 citations
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 260 citations
- Exploring Architectural Ingredients of Adversarially Robust Deep Neural NetworksHanxun Huang, Yisen Wang, Sarah M. Erfani, Quanquan Gu et al.NeurIPS 2021 · 124 citations
Related papers
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 169 citations
- Width Independent Bounds for the Local Lipschitz Constant of Deep Neural Networks at Random Initialization and after Lazy TrainingApostolos Evangelidis, Felix KrahmerICML 2026
- Benign Overfitting in Deep Neural Networks under Lazy TrainingZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Francesco Locatello et al.ICML 2023 · 12 citations
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 52 citations
- Subquadratic Overparameterization for Shallow Neural NetworksChaehwan Song, Ali Ramezani-Kebrya, Thomas Pethick, Armin Eftekhari et al.NeurIPS 2021 · 35 citations
