AdaLRS: Loss-Guided Adaptive Learning Rate Search for Efficient Foundation Model Pretraining
Hongyuan Dong, Dingkang Yang, Xiao Liang, Chao Feng, Ran Jiao
摘要
Learning rate is widely regarded as crucial for effective foundation model pretraining. Recent research explores and demonstrates the transferability of learning rate configurations across varying model and dataset sizes, etc. Nevertheless, these approaches are constrained to specific training scenarios and typically necessitate extensive hyperparameter tuning on proxy models. In this work, we propose AdaLRS, a plug-in-and-play adaptive learning rate search algorithm that conducts online optimal learning rate search via optimizing loss descent velocities. We provide theoretical and experimental analyzes to show that foundation model pretraining loss and its descent velocity are both convex and share the same optimal learning rate. Relying solely on training loss dynamics, AdaLRS involves few extra computations to guide the search process, and its convergence is guaranteed via theoretical analysis. Experiments on both LLM and VLM pretraining show that AdaLRS adjusts suboptimal learning rates to the neighborhood of optimum with marked efficiency and effectiveness, with model performance improved accordingly. We also show the robust generalizability of AdaLRS across varying training scenarios, such as different model sizes, training paradigms, base learning rate scheduler choices, and hyperparameter settings. 1 max(λ t β,1) , where α, β > 1 are two multiplicatively independent real numbers and λ ∈ (0, 1) is a decay factor. We validate the multiplicatively independent design of LR scaling factors (for all integers m, n, α m = β n =⇒ m = n = 0) in Appendix B.
Workflow. During model training, AdaLRS monitors the loss curve slope v t with the least squares method [7], and attempts to upscale the learning rate when the loss curve slope decays. After the upscaling adjustment, we compare the loss curve slope with that before upscaling. As shown in Equation 1, if the estimated loss slope increases more than 2e after the adjustment, the upscaling is regarded as valid and the adjustment will be retained. On the other hand, once the validation fails
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang 等NeurIPS 2024 · 被引用 1,029 次
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
相关 Paper
- AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the FlyYuchen Jin, Tianyi Zhou, Liangyu Zhao, Yibo Zhu 等ICLR 2021 · 被引用 26 次
- A Multi-Power Law for Loss Curve Prediction Across Learning Rate SchedulesKairong Luo, Haodong Wen, Shengding Hu, Zhenbo Sun 等ICLR 2025
- Learning-Rate-Free Learning by D-AdaptationAaron Defazio, Konstantin MishchenkoICML 2023 · 被引用 117 次
- Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-TuningKazuki Yano, Shun Kiyono, Sosuke Kobayashi, Sho Takase 等ICLR 2026 · 被引用 13 次
- Scaling and Transferability of Annealing Strategies in Large Language Model TrainingSiqi Wang, Zhengyu Chen, Teng Xiao, Zheqi Lv 等AAAI 2026 · 被引用 1 次
