The Flip Side of the Reweighted Coin: Duality of Adaptive Dropout and Regularization
Daniel LeJeune, Hamid Javadi, Richard G. Baraniuk
Abstract
Among the most successful methods for sparsifying deep (neural) networks are those that adaptively mask the network weights throughout training. By examining this masking, or dropout, in the linear case, we uncover a duality between such adaptive methods and regularization through the so-called"-trick"that casts both as iteratively reweighted optimizations. We show that any dropout strategy that adapts to the weights in a monotonic way corresponds to an effective subquadratic regularization penalty, and therefore leads to sparse solutions. We obtain the effective penalties for several popular sparsification strategies, which are remarkably similar to classical penalties commonly used in sparse optimization. Considering variational dropout as a case study, we demonstrate similar empirical behavior between the adaptive dropout method and classical methods on the task of deep network sparsification, validating our theory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18fff70a-40b4-40b4-8b30-cccc0a84eebdCited by top-tier papers3
- Asymptotically Free Sketched Ridge Ensembles: Risks, Cross-Validation, and TuningPratik Patil, Daniel LeJeuneICLR 2024 · 13 citations
- Implicit Regularization Paths of Weighted Neural RepresentationsJin-Hong Du, Pratik PatilNeurIPS 2024 · 2 citations
- MAST: model-agnostic sparsified trainingYury Demidovich, Grigory Malinovsky, Egor Shulgin, Peter RichtárikICLR 2025
Builds on3
- The Implicit and Explicit Regularization Effects of DropoutColin Wei, Sham M. Kakade, Tengyu MaICML 2020 · 129 citations
- Dropout: Explicit Forms and Capacity ControlRaman Arora, Peter L. Bartlett, Poorya Mianjy, Nathan SrebroICML 2021 · 43 citations
- On Convergence and Generalization of Dropout TrainingPoorya Mianjy, Raman AroraNeurIPS 2020 · 34 citations
Related papers
- On the Regularization Properties of Structured DropoutAmbar Pal, Connor Lane, René Vidal, Benjamin D. HaeffeleCVPR 2020
- Fiedler Regularization: Learning Neural Networks with Graph SparsityEdric Tam, David B. DunsonICML 2020 · 15 citations
- Structured Dropout Variational Inference for Bayesian Neural NetworksSon Nguyen, Duong Nguyen, Khai Nguyen, Khoat Than et al.NeurIPS 2021 · 11 citations
- Neural Pruning via Growing RegularizationHuan Wang, Can Qin, Yulun Zhang, Yun FuICLR 2021 · 188 citations
- On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm PerspectiveZeke Xie, Zhiqiang Xu, Jingzhao Zhang, Issei Sato et al.NeurIPS 2023 · 38 citations
