The Curse of Unrolling: Rate of Differentiating Through Optimization
Damien Scieur, Gauthier Gidel, Quentin Bertrand, Fabian Pedregosa
Abstract
Computing the Jacobian of the solution of an optimization problem is a central problem in machine learning, with applications in hyperparameter optimization, meta-learning, optimization as a layer, and dataset distillation, to name a few. Unrolled differentiation is a popular heuristic that approximates the solution using an iterative solver and differentiates it through the computational path. This work provides a non-asymptotic convergence-rate analysis of this approach on quadratic objectives for gradient descent and the Chebyshev method. We show that to ensure convergence of the Jacobian, we can either 1) choose a large learning rate leading to a fast asymptotic convergence but accept that the algorithm may have an arbitrarily long burn-in phase or 2) choose a smaller learning rate leading to an immediate but slower convergence. We refer to this phenomenon as the curse of unrolling. Finally, we discuss open problems relative to this approach, such as deriving a practical update rule for the optimal unrolling strategy and making novel connections with the field of Sobolev orthogonal polynomials.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- One-step differentiation of iterative algorithmsJérôme Bolte, Edouard Pauwels, Samuel VaiterNeurIPS 2023 · 36 citations
- FSNet: Feasibility-Seeking Neural Network for Constrained Optimization with GuaranteesHoang T. Nguyen, Priya L. DontiNeurIPS 2025 · 26 citations
- Let Go of Your Labels with Unsupervised TransferArtyom Gadetsky, Yulun Jiang, Maria BrbicICML 2024 · 16 citations
- Differentiation Through Black-Box Quadratic Programming SolversConnor W. Magoon, Fengyu Yang, Noam Aigerman, Shahar Z. KovalskyNeurIPS 2025 · 14 citations
- Leveraging augmented-Lagrangian techniques for differentiating over infeasible quadratic programs in machine learningAntoine Bambade, Fabian Schramm, Adrien B. Taylor, Justin CarpentierICLR 2024 · 9 citations
Builds on6
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig et al.NeurIPS 2022 · 386 citations
- On the Iteration Complexity of Hypergradient ComputationRiccardo Grazzi, Luca Franceschi, Massimiliano Pontil, Saverio SalzoICML 2020 · 241 citations
- Implicit differentiation of Lasso-type models for hyperparameter optimizationQuentin Bertrand, Quentin Klopfenstein, Mathieu Blondel, Samuel Vaiter et al.ICML 2020 · 73 citations
- Super-efficiency of automatic differentiation for functions defined as a minimumPierre Ablin, Gabriel Peyré, Thomas MoreauICML 2020 · 42 citations
- Acceleration via Fractal Learning Rate SchedulesNaman Agarwal, Surbhi Goel, Cyril ZhangICML 2021 · 19 citations
Related papers
- Online Hyperparameter Meta-Learning with Hypergradient DistillationHaebeom Lee, Hayeon Lee, Jaewoong Shin, Eunho Yang et al.ICLR 2022 · 6 citations
- Nonsmooth Implicit Differentiation: Deterministic and Stochastic Convergence RatesRiccardo Grazzi, Massimiliano Pontil, Saverio SalzoICML 2024 · 5 citations
- Bilevel Optimization: Convergence Analysis and Enhanced DesignKaiyi Ji, Junjie Yang, Yingbin LiangICML 2021 · 343 citations
- On Implicit Bias in Overparameterized Bilevel OptimizationPaul Vicol, Jonathan P. Lorraine, Fabian Pedregosa, David Duvenaud et al.ICML 2022 · 48 citations
- Understanding approximate and unrolled dictionary learning for pattern recoveryBenoît Malézieux, Thomas Moreau, Matthieu KowalskiICLR 2022 · 17 citations
