Where Do Large Learning Rates Lead Us?
Ildus Sadrtdinov, Maxim Kodryan, Eduard Pokonechny, Ekaterina Lobacheva, Dmitry P. Vetrov
摘要
It is generally accepted that starting neural networks training with large learning rates (LRs) improves generalization. Following a line of research devoted to understanding this effect, we conduct an empirical study in a controlled setting focusing on two questions: 1) how large an initial LR is required for obtaining optimal quality, and 2) what are the key differences between models trained with different LRs? We discover that only a narrow range of initial LRs slightly above the convergence threshold lead to optimal results after fine-tuning with a small LR or weight averaging. By studying the local geometry of reached minima, we observe that using LRs from this optimal range allows for the optimization to locate a basin that only contains high-quality minima. Additionally, we show that these initial LRs result in a sparse set of learned features, with a clear focus on those most relevant for the task. In contrast, starting training with too small LRs leads to unstable minima and attempts to learn all features simultaneously, resulting in poor generalization. Conversely, using initial LRs that are too large fails to detect a basin with good solutions and extract meaningful patterns from the data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMsJiacheng Lin, Zhongruo Wang, Kun Qian, Tian Wang 等ICLR 2026 · 被引用 25 次
- How does the optimizer implicitly bias the model merging loss landscape?Chenxiang Zhang, Alexander Theus, Damien Teney, Antonio Orvieto 等ICLR 2026 · 被引用 2 次
- Fixed Aggregation Features Can Rival GNNsCelia Rubio-Madrigal, Rebekka BurkholzICML 2026 · 被引用 2 次
它引用的顶会 Paper34
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 被引用 653 次
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
相关 Paper
- Benign Oscillation of Stochastic Gradient Descent with Large Learning RateMiao Lu, Beining Wu, Xiaodong Yang, Difan ZouICLR 2024 · 被引用 9 次
- Maximal Initial Learning Rates in Deep ReLU NetworksGaurav Iyer, Boris Hanin, David RolnickICML 2023 · 被引用 14 次
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit 等ICLR 2020 · 被引用 198 次
- How much does Initialization Affect Generalization?Sameera Ramasinghe, Lachlan Ewen MacDonald, Moshiur R. Farazi, Hemanth Saratchandran 等ICML 2023 · 被引用 9 次
- Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learningDaniel Kunin, Allan Raventós, Clémentine C. J. Dominé, Feng Chen 等NeurIPS 2024 · 被引用 48 次
