Regularization in ResNet with Stochastic Depth
Soufiane Hayou, Fadhel Ayed
Abstract
Regularization plays a major role in modern deep learning. From classic techniques such as L 1 , L 2 penalties to other noise-based methods such as Dropout, regularization often yields better generalization properties by avoiding overfitting. Recently, Stochastic Depth (SD) has emerged as an alternative regularization technique for residual neural networks (ResNets) and has proven to boost the performance of ResNet on many tasks [Huang et al., 2016] . Despite the recent success of SD, little is known about this technique from a theoretical perspective. This paper provides a hybrid analysis combining perturbation analysis and signal propagation to shed light on different regularization effects of SD. Our analysis allows us to derive principled guidelines for choosing the survival rates used for training with SD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
- Width and Depth Limits Commute in Residual NetworksSoufiane Hayou, Greg YangICML 2023 · 23 citations
Builds on3
- The Implicit and Explicit Regularization Effects of DropoutColin Wei, Sham M. Kakade, Tengyu MaICML 2020 · 129 citations
- Explicit Regularisation in Gaussian Noise InjectionsAlexander Camuto, Matthew Willetts, Umut Simsekli, Stephen J. Roberts et al.NeurIPS 2020 · 90 citations
- Robust Pruning at InitializationSoufiane Hayou, Jean-Francois Ton, Arnaud Doucet, Yee Whye TehICLR 2021 · 50 citations
Related papers
- How Does Noise Help Robustness? Explanation and Exploration under the Neural SDE FrameworkXuanqing Liu, Tesi Xiao, Si Si, Qin Cao et al.CVPR 2020
- Local Regularizer Improves GeneralizationYikai Zhang, Hui Qu, Dimitris N. Metaxas, Chao ChenAAAI 2020 · 4 citations
- Learning Efficient Image Super-Resolution Networks via Structure-Regularized PruningYulun Zhang, Huan Wang, Can Qin, Yun FuICLR 2022 · 61 citations
- Co-training 2L Submodels for Visual RecognitionHugo Touvron, Matthieu Cord, Maxime Oquab, Piotr Bojanowski et al.CVPR 2023
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 173 citations
