Robustness in deep learning: The good (width), the bad (depth), and the ugly (initialization)
Zhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Volkan Cevher
摘要
We study the average robustness notion in deep neural networks in (selected) wide and narrow, deep and shallow, as well as lazy and non-lazy training settings. We prove that in the under-parameterized setting, width has a negative effect while it improves robustness in the over-parameterized setting. The effect of depth closely depends on the initialization and the training mode. In particular, when initialized with LeCun initialization, depth helps robustness with the lazy training regime. In contrast, when initialized with Neural Tangent Kernel (NTK) and He-initialization, depth hurts the robustness. Moreover, under the non-lazy training regime, we demonstrate how the width of a two-layer ReLU network benefits robustness. Our theoretical developments improve the results by [Huang et al. NeurIPS21; Wu et al. NeurIPS21] and are consistent with [Bubeck and Sellke NeurIPS21; Bubeck et al. COLT21].
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Sample Complexity Bounds for Score-Matching: Causal Discovery and Generative ModelingZhenyu Zhu, Francesco Locatello, Volkan CevherNeurIPS 2023 · 被引用 23 次
- Initialization Matters: Privacy-Utility Analysis of Overparameterized Neural NetworksJiayuan Ye, Zhenyu Zhu, Fanghui Liu, Reza Shokri 等NeurIPS 2023 · 被引用 19 次
- GlucoBench: Curated List of Continuous Glucose Monitoring Datasets with Prediction BenchmarksRenat Sergazinov, Elizabeth Chun, Valeriya Rogovchenko, Nathaniel J. Fernandes 等ICLR 2024 · 被引用 13 次
- Beyond the Universal Law of Robustness: Sharper Laws for Random Features and Neural Tangent KernelsSimone Bombari, Shayan Kiyani, Marco MondelliICML 2023 · 被引用 13 次
- Robust NAS under adversarial training: benchmark, theory, and beyondYongtao Wu, Fanghui Liu, Carl-Johann Simon-Gabriel, Grigorios Chrysos 等ICLR 2024 · 被引用 10 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Proving the Lottery Ticket Hypothesis: Pruning is All You NeedEran Malach, Gilad Yehudai, Shai Shalev-Shwartz, Ohad ShamirICML 2020 · 被引用 327 次
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 被引用 260 次
- Exploring Architectural Ingredients of Adversarially Robust Deep Neural NetworksHanxun Huang, Yisen Wang, Sarah M. Erfani, Quanquan Gu 等NeurIPS 2021 · 被引用 124 次
相关 Paper
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 被引用 169 次
- Width Independent Bounds for the Local Lipschitz Constant of Deep Neural Networks at Random Initialization and after Lazy TrainingApostolos Evangelidis, Felix KrahmerICML 2026
- Benign Overfitting in Deep Neural Networks under Lazy TrainingZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Francesco Locatello 等ICML 2023 · 被引用 12 次
- On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear WidthsQuynh NguyenICML 2021 · 被引用 52 次
- Subquadratic Overparameterization for Shallow Neural NetworksChaehwan Song, Ali Ramezani-Kebrya, Thomas Pethick, Armin Eftekhari 等NeurIPS 2021 · 被引用 35 次
