Parabolic Approximation Line Search for DNNs
Maximus Mutschler, Andreas Zell
Abstract
A major challenge in current optimization research for deep learning is automatically finding optimal step sizes for each update step. The optimal step size is closely related to the shape of the loss in the update step direction. However, this shape has not yet been examined in detail. This work shows empirically that the mini-batch loss along lines in negative gradient direction is locally mostly convex and well suited for one-dimensional parabolic approximations. We introduce a simple and robust line search approach by exploiting this parabolic observation, which performs loss-shape-dependent update steps. Our approach combines well-known methods such as parabolic approximation, line search, and conjugate gradient to perform efficiently. It surpasses other step size estimating methods and competes with standard optimization methods on a large variety of experiments without the need for hand-designed step size schedules. Thus, it is of interest for objectives where step-size schedules are unknown or do not perform well. Our extensive evaluation includes multiple comprehensive hyperparameter grid searches on several datasets and architectures. Finally, we provide a general investigation of exact line searches in the context of batch losses and exact losses, including their relation to our line search approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f059e994-92a2-4d75-bbc6-7b764b0734afCited by top-tier papers3
- Descending through a Crowded Valley - Benchmarking Deep Learning OptimizersRobin M. Schmidt, Frank Schneider, Philipp HennigICML 2021 · 195 citations
- Don't be so Monotone: Relaxing Stochastic Line Search in Over-Parameterized ModelsLeonardo Galli, Holger Rauhut, Mark SchmidtNeurIPS 2023 · 20 citations
- Stepping on the Edge: Curvature Aware Learning Rate TunersVincent Roulet, Atish Agarwala, Jean-Bastien Grill, Grzegorz Swirszcz et al.NeurIPS 2024 · 9 citations
Builds on2
Related papers
- QLABGrad: A Hyperparameter-Free and Convergence-Guaranteed Scheme for Deep LearningMinghan Fu, Fang-Xiang WuAAAI 2024 · 12 citations
- A Second look at Exponential and Cosine Step Sizes: Simplicity, Adaptivity, and PerformanceXiaoyu Li, Zhenxun Zhuang, Francesco OrabonaICML 2021 · 29 citations
- Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning RateZhiqi Bu, Shiyun Xu, Jialin MaoICLR 2026 · 4 citations
- On the Convergence of Step Decay Step-Size for Stochastic OptimizationXiaoyu Wang, Sindri Magnússon, Mikael JohanssonNeurIPS 2021 · 33 citations
- An Even More Optimal Stochastic Optimization Algorithm: Minibatching and Interpolation LearningBlake E. Woodworth, Nathan SrebroNeurIPS 2021 · 22 citations
