The Implicit and Explicit Regularization Effects of Dropout
Colin Wei, Sham M. Kakade, Tengyu Ma
Abstract
Dropout is a widely-used regularization technique, often required to obtain state-of-the-art for a number of architectures. This work demonstrates that dropout introduces two distinct but entangled regularization effects: an explicit effect (also studied in prior work) which occurs since dropout modifies the expected training objective, and, perhaps surprisingly, an additional implicit effect from the stochasticity in the dropout training update. This implicit regularization effect is analogous to the effect of stochasticity in small mini-batch stochastic gradient descent. We disentangle these two effects through controlled experiments. We then derive analytic simplifications which characterize each effect in terms of the derivatives of the model and the loss, for deep neural networks. We demonstrate these simplified, analytic regularizers accurately capture the important aspects of dropout, showing they faithfully replace dropout in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdef3128-a348-49c1-bd33-7bbad701dd31Cited by top-tier papers41
- Understanding the Role of Training Regimes in Continual LearningSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Razvan Pascanu, Hassan GhasemzadehNeurIPS 2020 · 295 citations
- How Does Mixup Help With Robustness and Generalization?Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani et al.ICLR 2021 · 294 citations
- Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-trainingHong Liu, Zhiyuan Li, David Leo Wright Hall, Percy Liang et al.ICLR 2024 · 264 citations
- Implicit Regularization in Deep Learning May Not Be Explainable by NormsNoam Razin, Nadav CohenNeurIPS 2020 · 178 citations
- Vision Transformers provably learn spatial structureSamy Jelassi, Michael E. Sander, Yuanzhi LiNeurIPS 2022 · 115 citations
Related papers
- On the Regularization Properties of Structured DropoutAmbar Pal, Connor Lane, René Vidal, Benjamin D. HaeffeleCVPR 2020
- Dropout Reduces UnderfittingZhuang Liu, Zhiqiu Xu, Joseph Jin, Zhiqiang Shen et al.ICML 2023 · 60 citations
- Deconstructing the Regularization of BatchNormYann N. Dauphin, Ekin Dogus CubukICLR 2021 · 6 citations
- Stochastic Modified Equations and Dynamics of Dropout AlgorithmZhongwang Zhang, Yuqing Li, Tao Luo, Zhi-Qin John XuICLR 2024 · 12 citations
- Implicit Jacobian regularization weighted with impurity of probability outputSungyoon Lee, Jinseong Park, Jaewook LeeICML 2023 · 6 citations
