Automatic Differentiation of Optimization Algorithms with Time-Varying Updates
Sheheryar Mehmood, Peter Ochs
摘要
Numerous Optimization Algorithms have a time-varying update rule thanks to, for instance, a changing step size, momentum parameter or, Hessian approximation. In this paper, we apply unrolled or automatic differentiation to a time-varying iterative process and provide convergence (rate) guarantees for the resulting derivative iterates. We adapt these convergence results and apply them to proximal gradient descent with variable step size and FISTA when solving partly smooth problems. We confirm our findings numerically by solving ℓ 1 and ℓ 2 -regularized linear and logisitc regression respectively. Our theoretical and numerical results show that the convergence rate of the algorithm is reflected in its derivative iterates. x∈X F (x, u) , (P) with a smooth objective F : X ×U → R. In this case (R e ) reduces to the optimality condition ∇ x F (x, u) = 0.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig 等NeurIPS 2022 · 被引用 386 次
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 被引用 272 次
- Monotone operator equilibrium networksEzra Winston, J. Zico KolterNeurIPS 2020 · 被引用 177 次
- BOME! Bilevel Optimization Made Easy: A Simple First-Order ApproachBo Liu, Mao Ye, Stephen Wright, Peter Stone 等NeurIPS 2022 · 被引用 170 次
- A Fully First-Order Method for Stochastic Bilevel OptimizationJeongyeol Kwon, Dohyun Kwon, Stephen Wright, Robert D. NowakICML 2023 · 被引用 123 次
相关 Paper
- Learning to solve TV regularised problems with unrolled algorithmsHamza Cherkaoui, Jeremias Sulam, Thomas MoreauNeurIPS 2020 · 被引用 16 次
- Directional Smoothness and Gradient Methods: Convergence and AdaptivityAaron Mishkin, Ahmed Khaled, Yuanhao Wang, Aaron Defazio 等NeurIPS 2024 · 被引用 25 次
- Estimating Generalization Performance Along the Trajectory of Proximal SGD in Robust RegressionKai Tan, Pierre C. BellecNeurIPS 2024
- Stochastic optimization under time drift: iterate averaging, step-decay schedules, and high probability guaranteesJoshua Cutler, Dmitriy Drusvyatskiy, Zaïd HarchaouiNeurIPS 2021 · 被引用 19 次
- The Curse of Unrolling: Rate of Differentiating Through OptimizationDamien Scieur, Gauthier Gidel, Quentin Bertrand, Fabian PedregosaNeurIPS 2022 · 被引用 20 次
