On the Regularization Properties of Structured Dropout
Ambar Pal, Connor Lane, René Vidal, Benjamin D. Haeffele
Abstract
Dropout and its extensions (e.g. DropBlock and Drop-Connect) are popular heuristics for training neural networks, which have been shown to improve generalization performance in practice. However, a theoretical understanding of their optimization and regularization properties remains elusive. Recent work shows that in the case of single hidden-layer linear networks, Dropout is a stochastic gradient descent method for minimizing a regularized loss, and that the regularizer induces solutions that are lowrank and balanced. In this work we show that for single hidden-layer linear networks, DropBlock induces spectral k-support norm regularization, and promotes solutions that are low-rank and have factors with equal norm. We also show that the global minimizer for DropBlock can be computed in closed form, and that DropConnect is equivalent to Dropout. We then show that some of these results can be extended to a general class of Dropout-strategies, and, with some assumptions, to deep non-linear networks when Dropout is applied to the last layer. We verify our theoretical claims and assumptions experimentally with commonly used network architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 760a9779-6fbc-45e5-8810-2f9e407cccd8Cited by top-tier papers1
Ask how each one uses itRelated papers
- The Implicit and Explicit Regularization Effects of DropoutColin Wei, Sham M. Kakade, Tengyu MaICML 2020 · 129 citations
- The Flip Side of the Reweighted Coin: Duality of Adaptive Dropout and RegularizationDaniel LeJeune, Hamid Javadi, Richard G. BaraniukNeurIPS 2021 · 8 citations
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural NetworksAlexander Shevchenko, Marco MondelliICML 2020 · 41 citations
- Batch normalization provably avoids ranks collapse for randomly initialised deep networksHadi Daneshmand, Jonas Moritz Kohler, Francis R. Bach, Thomas Hofmann et al.NeurIPS 2020 · 73 citations
- Stochastic Modified Equations and Dynamics of Dropout AlgorithmZhongwang Zhang, Yuqing Li, Tao Luo, Zhi-Qin John XuICLR 2024 · 12 citations
