Searching for Optimal Per-Coordinate Step-sizes with Multidimensional Backtracking
Frederik Kunstner, Victor Sanches Portella, Mark Schmidt, Nicholas J. A. Harvey
摘要
The backtracking line-search is an effective technique to automatically tune the step-size in smooth optimization. It guarantees similar performance to using the theoretically optimal step-size. Many approaches have been developed to instead tune per-coordinate step-sizes, also known as diagonal preconditioners, but none of the existing methods are provably competitive with the optimal per-coordinate stepsizes. We propose multidimensional backtracking, an extension of the backtracking line-search to find good diagonal preconditioners for smooth convex problems. Our key insight is that the gradient with respect to the step-sizes, also known as hypergradients, yields separating hyperplanes that let us search for good preconditioners using cutting-plane methods. As black-box cutting-plane approaches like the ellipsoid method are computationally prohibitive, we develop an efficient algorithm tailored to our setting. Multidimensional backtracking is provably competitive with the best diagonal preconditioner and requires no manual tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Adaptive SGD with Polyak stepsize and Line-search: Robust Convergence and Variance ReductionXiaowen Jiang, Sebastian U. StichNeurIPS 2023 · 被引用 40 次
- Provable and Practical Online Learning Rate Adaptation with Hypergradient DescentYa-Chi Chu, Wenzhi Gao, Yinyu Ye, Madeleine UdellICML 2025
- MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parametersArsalan Sharifnassab, Saber Salehkaleybar, Richard S. SuttonICML 2025
它引用的顶会 Paper5
- ADAHESSIAN: An Adaptive Second Order Optimizer for Machine LearningZhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa 等AAAI 2021 · 被引用 358 次
- Gradient Descent: The Ultimate OptimizerKartik Chandra, Audrey Xie, Jonathan Ragan-Kelley, Erik MeijerNeurIPS 2022 · 被引用 66 次
- Smoothness Matrices Beat Smoothness Constants: Better Communication Compression Techniques for Distributed OptimizationMher Safaryan, Filip Hanzely, Peter RichtárikNeurIPS 2021 · 被引用 32 次
- Doubly Adaptive Scaled Algorithm for Machine Learning Using Second-Order InformationMajid Jahani, Sergey Rusakov, Zheng Shi, Peter Richtárik 等ICLR 2022 · 被引用 31 次
- Amortized Proximal OptimizationJuhan Bae, Paul Vicol, Jeff Z. HaoChen, Roger B. GrosseNeurIPS 2022 · 被引用 15 次
相关 Paper
- Adaptive backtracking line searchJoao V. Cavalcanti, Laurent Lessard, Ashia C. WilsonICLR 2025
- Parabolic Approximation Line Search for DNNsMaximus Mutschler, Andreas ZellNeurIPS 2020 · 被引用 22 次
- Newton Method Revisited: Global Convergence Rates up to O(1/k3) for Stepsize Schedules and Linesearch ProceduresSlavomír Hanzely, Farshed Abdukhakimov, Martin TakácICLR 2026
- Achieving Linear Convergence with Parameter-Free Algorithms in Decentralized OptimizationIlya A. Kuruzov, Gesualdo Scutari, Alexander V. GasnikovNeurIPS 2024 · 被引用 10 次
- Polynomial Preconditioning for Gradient MethodsNikita Doikov, Anton RodomanovICML 2023 · 被引用 2 次
