The Flip Side of the Reweighted Coin: Duality of Adaptive Dropout and Regularization
Daniel LeJeune, Hamid Javadi, Richard G. Baraniuk
摘要
Among the most successful methods for sparsifying deep (neural) networks are those that adaptively mask the network weights throughout training. By examining this masking, or dropout, in the linear case, we uncover a duality between such adaptive methods and regularization through the so-called"-trick"that casts both as iteratively reweighted optimizations. We show that any dropout strategy that adapts to the weights in a monotonic way corresponds to an effective subquadratic regularization penalty, and therefore leads to sparse solutions. We obtain the effective penalties for several popular sparsification strategies, which are remarkably similar to classical penalties commonly used in sparse optimization. Considering variational dropout as a case study, we demonstrate similar empirical behavior between the adaptive dropout method and classical methods on the task of deep network sparsification, validating our theory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Asymptotically Free Sketched Ridge Ensembles: Risks, Cross-Validation, and TuningPratik Patil, Daniel LeJeuneICLR 2024 · 被引用 13 次
- Implicit Regularization Paths of Weighted Neural RepresentationsJin-Hong Du, Pratik PatilNeurIPS 2024 · 被引用 2 次
- MAST: model-agnostic sparsified trainingYury Demidovich, Grigory Malinovsky, Egor Shulgin, Peter RichtárikICLR 2025
它引用的顶会 Paper3
- The Implicit and Explicit Regularization Effects of DropoutColin Wei, Sham M. Kakade, Tengyu MaICML 2020 · 被引用 129 次
- Dropout: Explicit Forms and Capacity ControlRaman Arora, Peter L. Bartlett, Poorya Mianjy, Nathan SrebroICML 2021 · 被引用 43 次
- On Convergence and Generalization of Dropout TrainingPoorya Mianjy, Raman AroraNeurIPS 2020 · 被引用 34 次
相关 Paper
- On the Regularization Properties of Structured DropoutAmbar Pal, Connor Lane, René Vidal, Benjamin D. HaeffeleCVPR 2020
- Fiedler Regularization: Learning Neural Networks with Graph SparsityEdric Tam, David B. DunsonICML 2020 · 被引用 15 次
- Structured Dropout Variational Inference for Bayesian Neural NetworksSon Nguyen, Duong Nguyen, Khai Nguyen, Khoat Than 等NeurIPS 2021 · 被引用 11 次
- Neural Pruning via Growing RegularizationHuan Wang, Can Qin, Yulun Zhang, Yun FuICLR 2021 · 被引用 188 次
- On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm PerspectiveZeke Xie, Zhiqiang Xu, Jingzhao Zhang, Issei Sato 等NeurIPS 2023 · 被引用 38 次
