Reparameterizing Mirror Descent as Gradient Descent
Ehsan Amid, Manfred K. Warmuth
Abstract
Most of the recent successful applications of neural networks have been based on training with gradient descent updates. However, for some small networks, other mirror descent updates learn provably more efficiently when the target is sparse. We present a general framework for casting a mirror descent update as a gradient descent update on a different set of parameters. In some cases, the mirror descent reparameterization can be described as training a modified network with standard backpropagation. The reparameterization framework is versatile and covers a wide range of mirror descent updates, even cases where the domain is constrained. Our construction for the reparameterization argument is done for the continuous versions of the updates. Finding general criteria for the discrete versions to closely track their continuous counterparts remains an interesting open problem. ⇤ An earlier version of this manuscript (with additional results on the matrix case) appeared as "Interpolating Between Gradient Descent and Exponentiated Gradient Using Reparameterized Gradient Descent" as a preprint. 2 The normalized version is called EG and the two-sided version EGU ± . More about this later. 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fcf670af-42f8-405a-bc4d-26b6363c1ac2Cited by top-tier papers10
- On the Implicit Bias of Initialization Shape: Beyond Infinitesimal Mirror DescentShahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E. Woodworth et al.ICML 2021 · 85 citations
- Max-Margin Token Selection in Attention MechanismDavoud Ataee Tarzanagh, Yingcong Li, Xuechen Zhang, Samet OymakNeurIPS 2023 · 67 citations
- Implicit Bias of Gradient Descent on Reparametrized Models: On Equivalence to Mirror DescentZhiyuan Li, Tianhao Wang, Jason D. Lee, Sanjeev AroraNeurIPS 2022 · 49 citations
- Non-convex online learning via algorithmic equivalenceUdaya Ghai, Zhou Lu, Elad HazanNeurIPS 2022 · 15 citations
- Identifying Equivalent Training DynamicsWilliam T. Redman, Juan M. Bello-Rivas, Maria Fonoberova, Ryan Mohr et al.NeurIPS 2024 · 15 citations
Related papers
- Implicit Regularization for Group SparsityJiangyuan Li, Thanh Van Nguyen, Chinmay Hegde, Raymond K. W. WongICLR 2023 · 2 citations
- Never Saddle for Reparameterized Steepest Descent as Mirror FlowTom Jacobs, Chao Zhou, Rebekka BurkholzICLR 2026 · 3 citations
- Mirror Descent Maximizes Generalized Margin and Can Be Implemented EfficientlyHaoyuan Sun, Kwangjun Ahn, Christos Thrampoulidis, Navid AzizanNeurIPS 2022 · 33 citations
- A Novel Framework for Policy Mirror Descent with General Parameterization and Linear ConvergenceCarlo Alfano, Rui Yuan, Patrick RebeschiniNeurIPS 2023 · 25 citations
- Meta-Learning with Warped Gradient DescentSebastian Flennerhag, Andrei A. Rusu, Razvan Pascanu, Francesco Visin et al.ICLR 2020 · 221 citations
