On Warm-Starting Neural Network Training
Jordan T. Ash, Ryan P. Adams
Abstract
In many real-world deployments of machine learning systems, data arrive piecemeal. These learning scenarios may be passive, where data arrive incrementally due to structural properties of the problem (e.g., daily financial data) or active, where samples are selected according to a measure of their quality (e.g., experimental design). In both of these cases, we are building a sequence of models that incorporate an increasing amount of data. We would like each of these models in the sequence to be performant and take advantage of all the data that are available to that point. Conventional intuition suggests that when solving a sequence of related optimization problems of this form, it should be possible to initialize using the solution of the previous iterate -- to "warm start" the optimization rather than initialize from scratch -- and see reductions in wall-clock time. However, in practice this warm-starting seems to yield poorer generalization performance than models that have fresh random initializations, even though the final training losses are similar. While it appears that some hyperparameter settings allow a practitioner to close this generalization gap, they seem to only do so in regimes that damage the wall-clock gains of the warm start. Nevertheless, it is highly desirable to be able to warm-start neural network training, as it would dramatically reduce the resource usage associated with the construction of performant deep learning systems. In this work, we take a closer look at this empirical phenomenon and try to understand when and how it occurs. We also provide a surprisingly simple trick that overcomes this pathology in several important situations, and present experiments that elucidate some of its properties.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers78
- Zeus: Understanding and Optimizing GPU Energy Consumption of DNN TrainingJie You, Jae-Won Chung, Mosharaf ChowdhuryNSDI 2023 · 220 citations
- If Influence Functions are the Answer, Then What is the Question?Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi et al.NeurIPS 2022 · 185 citations
- Understanding Plasticity in Neural NetworksClare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires et al.ICML 2023 · 162 citations
- The Dormant Neuron Phenomenon in Deep Reinforcement LearningGhada Sokar, Rishabh Agarwal, Pablo Samuel Castro, Utku EvciICML 2023 · 153 citations
- Continual World: A Robotic Benchmark For Continual Reinforcement LearningMaciej Wolczyk, Michal Zajac, Razvan Pascanu, Lukasz Kucinski et al.NeurIPS 2021 · 152 citations
Builds on2
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
Related papers
- Why Warmup the Learning Rate? Underlying Mechanisms and ImprovementsDayal Singh Kalra, Maissam BarkeshliNeurIPS 2024 · 87 citations
- DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of PlasticityBaekrok Shin, Junsoo Oh, Hanseul Cho, Chulhee YunNeurIPS 2024 · 8 citations
- ModelKeeper: Accelerating DNN Training via Automated Training WarmupFan Lai, Yinwei Dai, Harsha V. Madhyastha, Mosharaf ChowdhuryNSDI 2023 · 31 citations
- Where Do Large Learning Rates Lead Us?Ildus Sadrtdinov, Maxim Kodryan, Eduard Pokonechny, Ekaterina Lobacheva et al.NeurIPS 2024 · 6 citations
- How much does Initialization Affect Generalization?Sameera Ramasinghe, Lachlan Ewen MacDonald, Moshiur R. Farazi, Hemanth Saratchandran et al.ICML 2023 · 9 citations
