Parabolic Approximation Line Search for DNNs
Maximus Mutschler, Andreas Zell
摘要
A major challenge in current optimization research for deep learning is automatically finding optimal step sizes for each update step. The optimal step size is closely related to the shape of the loss in the update step direction. However, this shape has not yet been examined in detail. This work shows empirically that the mini-batch loss along lines in negative gradient direction is locally mostly convex and well suited for one-dimensional parabolic approximations. We introduce a simple and robust line search approach by exploiting this parabolic observation, which performs loss-shape-dependent update steps. Our approach combines well-known methods such as parabolic approximation, line search, and conjugate gradient to perform efficiently. It surpasses other step size estimating methods and competes with standard optimization methods on a large variety of experiments without the need for hand-designed step size schedules. Thus, it is of interest for objectives where step-size schedules are unknown or do not perform well. Our extensive evaluation includes multiple comprehensive hyperparameter grid searches on several datasets and architectures. Finally, we provide a general investigation of exact line searches in the context of batch losses and exact losses, including their relation to our line search approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Descending through a Crowded Valley - Benchmarking Deep Learning OptimizersRobin M. Schmidt, Frank Schneider, Philipp HennigICML 2021 · 被引用 195 次
- Don't be so Monotone: Relaxing Stochastic Line Search in Over-Parameterized ModelsLeonardo Galli, Holger Rauhut, Mark SchmidtNeurIPS 2023 · 被引用 20 次
- Stepping on the Edge: Curvature Aware Learning Rate TunersVincent Roulet, Atish Agarwala, Jean-Bastien Grill, Grzegorz Swirszcz 等NeurIPS 2024 · 被引用 9 次
它引用的顶会 Paper2
相关 Paper
- QLABGrad: A Hyperparameter-Free and Convergence-Guaranteed Scheme for Deep LearningMinghan Fu, Fang-Xiang WuAAAI 2024 · 被引用 12 次
- A Second look at Exponential and Cosine Step Sizes: Simplicity, Adaptivity, and PerformanceXiaoyu Li, Zhenxun Zhuang, Francesco OrabonaICML 2021 · 被引用 29 次
- Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning RateZhiqi Bu, Shiyun Xu, Jialin MaoICLR 2026 · 被引用 4 次
- On the Convergence of Step Decay Step-Size for Stochastic OptimizationXiaoyu Wang, Sindri Magnússon, Mikael JohanssonNeurIPS 2021 · 被引用 33 次
- An Even More Optimal Stochastic Optimization Algorithm: Minibatching and Interpolation LearningBlake E. Woodworth, Nathan SrebroNeurIPS 2021 · 被引用 22 次
