QLABGrad: A Hyperparameter-Free and Convergence-Guaranteed Scheme for Deep Learning
Minghan Fu, Fang-Xiang Wu
Abstract
The learning rate is a critical hyperparameter for deep learning tasks since it determines the extent to which the model parameters are adjusted during the learning course. However, the choice of learning rates typically depends on empirical judgment, which may not result in satisfactory outcomes without intensive try-and-error experiments. In this study, we propose a novel learning rate adaptation scheme called QLABGrad. Without any user-specified hyperparameter, QLABGrad automatically determines the learning rate by optimizing the quadratic loss approximation-based (QLAB) function for a given gradient descent direction, where only one extra forward propagation is required. We theoretically prove the convergence of QLABGrad under the smooth Lipschitz condition on the loss function. Experiment results on multiple architectures, including MLP, CNN, and ResNet, on MNIST, CIFAR10, and ImageNet datasets, demonstrate that QLABGrad outperforms widely adopted schemes for deep learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 349800df-2a48-4331-af6a-f4f998f3e8c8Cited by top-tier papers1
Ask how each one uses itRelated papers
- Learning-Rate-Free Learning by D-AdaptationAaron Defazio, Konstantin MishchenkoICML 2023 · 117 citations
- Parabolic Approximation Line Search for DNNsMaximus Mutschler, Andreas ZellNeurIPS 2020 · 22 citations
- AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the FlyYuchen Jin, Tianyi Zhou, Liangyu Zhao, Yibo Zhu et al.ICLR 2021 · 26 citations
- Mechanic: A Learning Rate TunerAshok Cutkosky, Aaron Defazio, Harsh MehtaNeurIPS 2023 · 27 citations
- Large Batch Optimization for Deep Learning Using New Complete Layer-Wise Adaptive Rate ScalingZhouyuan Huo, Bin Gu, Heng HuangAAAI 2021 · 35 citations
