The Curse of Unrolling: Rate of Differentiating Through Optimization
Damien Scieur, Gauthier Gidel, Quentin Bertrand, Fabian Pedregosa
摘要
Computing the Jacobian of the solution of an optimization problem is a central problem in machine learning, with applications in hyperparameter optimization, meta-learning, optimization as a layer, and dataset distillation, to name a few. Unrolled differentiation is a popular heuristic that approximates the solution using an iterative solver and differentiates it through the computational path. This work provides a non-asymptotic convergence-rate analysis of this approach on quadratic objectives for gradient descent and the Chebyshev method. We show that to ensure convergence of the Jacobian, we can either 1) choose a large learning rate leading to a fast asymptotic convergence but accept that the algorithm may have an arbitrarily long burn-in phase or 2) choose a smaller learning rate leading to an immediate but slower convergence. We refer to this phenomenon as the curse of unrolling. Finally, we discuss open problems relative to this approach, such as deriving a practical update rule for the optimal unrolling strategy and making novel connections with the field of Sobolev orthogonal polynomials.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- One-step differentiation of iterative algorithmsJérôme Bolte, Edouard Pauwels, Samuel VaiterNeurIPS 2023 · 被引用 36 次
- FSNet: Feasibility-Seeking Neural Network for Constrained Optimization with GuaranteesHoang T. Nguyen, Priya L. DontiNeurIPS 2025 · 被引用 26 次
- Let Go of Your Labels with Unsupervised TransferArtyom Gadetsky, Yulun Jiang, Maria BrbicICML 2024 · 被引用 16 次
- Differentiation Through Black-Box Quadratic Programming SolversConnor W. Magoon, Fengyu Yang, Noam Aigerman, Shahar Z. KovalskyNeurIPS 2025 · 被引用 14 次
- Leveraging augmented-Lagrangian techniques for differentiating over infeasible quadratic programs in machine learningAntoine Bambade, Fabian Schramm, Adrien B. Taylor, Justin CarpentierICLR 2024 · 被引用 9 次
它引用的顶会 Paper6
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig 等NeurIPS 2022 · 被引用 386 次
- On the Iteration Complexity of Hypergradient ComputationRiccardo Grazzi, Luca Franceschi, Massimiliano Pontil, Saverio SalzoICML 2020 · 被引用 241 次
- Implicit differentiation of Lasso-type models for hyperparameter optimizationQuentin Bertrand, Quentin Klopfenstein, Mathieu Blondel, Samuel Vaiter 等ICML 2020 · 被引用 73 次
- Super-efficiency of automatic differentiation for functions defined as a minimumPierre Ablin, Gabriel Peyré, Thomas MoreauICML 2020 · 被引用 42 次
- Acceleration via Fractal Learning Rate SchedulesNaman Agarwal, Surbhi Goel, Cyril ZhangICML 2021 · 被引用 19 次
相关 Paper
- Online Hyperparameter Meta-Learning with Hypergradient DistillationHaebeom Lee, Hayeon Lee, Jaewoong Shin, Eunho Yang 等ICLR 2022 · 被引用 6 次
- Nonsmooth Implicit Differentiation: Deterministic and Stochastic Convergence RatesRiccardo Grazzi, Massimiliano Pontil, Saverio SalzoICML 2024 · 被引用 5 次
- Bilevel Optimization: Convergence Analysis and Enhanced DesignKaiyi Ji, Junjie Yang, Yingbin LiangICML 2021 · 被引用 343 次
- On Implicit Bias in Overparameterized Bilevel OptimizationPaul Vicol, Jonathan P. Lorraine, Fabian Pedregosa, David Duvenaud 等ICML 2022 · 被引用 48 次
- Understanding approximate and unrolled dictionary learning for pattern recoveryBenoît Malézieux, Thomas Moreau, Matthieu KowalskiICLR 2022 · 被引用 17 次
