Automatic Differentiation of Optimization Algorithms with Time-Varying Updates
Sheheryar Mehmood, Peter Ochs
Abstract
Numerous Optimization Algorithms have a time-varying update rule thanks to, for instance, a changing step size, momentum parameter or, Hessian approximation. In this paper, we apply unrolled or automatic differentiation to a time-varying iterative process and provide convergence (rate) guarantees for the resulting derivative iterates. We adapt these convergence results and apply them to proximal gradient descent with variable step size and FISTA when solving partly smooth problems. We confirm our findings numerically by solving ℓ 1 and ℓ 2 -regularized linear and logisitc regression respectively. Our theoretical and numerical results show that the convergence rate of the algorithm is reflected in its derivative iterates. x∈X F (x, u) , (P) with a smooth objective F : X ×U → R. In this case (R e ) reduces to the optimality condition ∇ x F (x, u) = 0.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig et al.NeurIPS 2022 · 386 citations
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 272 citations
- Monotone operator equilibrium networksEzra Winston, J. Zico KolterNeurIPS 2020 · 177 citations
- BOME! Bilevel Optimization Made Easy: A Simple First-Order ApproachBo Liu, Mao Ye, Stephen Wright, Peter Stone et al.NeurIPS 2022 · 170 citations
- A Fully First-Order Method for Stochastic Bilevel OptimizationJeongyeol Kwon, Dohyun Kwon, Stephen Wright, Robert D. NowakICML 2023 · 123 citations
Related papers
- Learning to solve TV regularised problems with unrolled algorithmsHamza Cherkaoui, Jeremias Sulam, Thomas MoreauNeurIPS 2020 · 16 citations
- Directional Smoothness and Gradient Methods: Convergence and AdaptivityAaron Mishkin, Ahmed Khaled, Yuanhao Wang, Aaron Defazio et al.NeurIPS 2024 · 25 citations
- Estimating Generalization Performance Along the Trajectory of Proximal SGD in Robust RegressionKai Tan, Pierre C. BellecNeurIPS 2024
- Stochastic optimization under time drift: iterate averaging, step-decay schedules, and high probability guaranteesJoshua Cutler, Dmitriy Drusvyatskiy, Zaïd HarchaouiNeurIPS 2021 · 19 citations
- The Curse of Unrolling: Rate of Differentiating Through OptimizationDamien Scieur, Gauthier Gidel, Quentin Bertrand, Fabian PedregosaNeurIPS 2022 · 20 citations
