ModelKeeper: Accelerating DNN Training via Automated Training Warmup
Fan Lai, Yinwei Dai, Harsha V. Madhyastha, Mosharaf Chowdhury
Abstract
With growing deployment of machine learning (ML) models, ML developers are training or re-training increasingly more deep neural networks (DNNs). They do so to find the most suitable model that meets their accuracy requirement while satisfying the resource and timeliness constraints of the target environment. In large shared clusters, the growing number of neural architecture search (NAS) and training jobs often result in models sharing architectural similarities with others from the same or a different ML developer. However, existing solutions do not provide a systematic mechanism to identify and leverage such similarities.
We present ModelKeeper, the first automated training warmup system that accelerates DNN training by repurposing previously-trained models in a shared cluster. Our key insight is that initializing a training job's model by transforming an already-trained model's weights can jump-start it and reduce the total amount of training needed. However, models submitted over time can differ in their architectures and accuracy. Given a new model to train, ModelKeeper scalably identifies its architectural similarity with previously trained models, selects a parent model with high similarity and good model accuracy, and performs structure-aware transformation of weights to preserve maximal information from the parent model during the warmup of new model weights. Our evaluations across thousands of CV and NLP models show that ModelKeeper achieves 1.3×-4.3× faster training completion with little overhead and no reduction in model accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 39b5d35c-7249-42aa-8f96-ac42544199d1Cited by top-tier papers8
- Approximate Caching for Efficiently Serving Text-to-Image Diffusion ModelsShubham Agarwal, Subrata Mitra, Sarthak Chakraborty, Srikrishna Karanam et al.NSDI 2024 · 44 citations
- USHER: Holistic Interference Avoidance for Resource Optimized ML InferenceSudipta Saha Shubha, Haiying Shen, Anand P. IyerOSDI 2024 · 35 citations
- Optimus: Warming Serverless ML Inference via Inter-Function Model TransformationZicong Hong, Jian Lin, Song Guo, Sifu Luo et al.EuroSys 2024 · 29 citations
- AdaEmbed: Adaptive Embedding for Large-Scale Recommendation ModelsFan Lai, Wei Zhang, Rui Liu, William Tsai et al.OSDI 2023 · 23 citations
- Hydro: Surrogate-Based Hyperparameter Tuning Service in DatacentersQinghao Hu, Zhisheng Ye, Meng Zhang, Qiaoling Chen et al.OSDI 2023 · 16 citations
Builds on18
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 825 citations
- What is being transferred in transfer learning?Behnam Neyshabur, Hanie Sedghi, Chiyuan ZhangNeurIPS 2020 · 654 citations
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao et al.OSDI 2020 · 392 citations
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
Related papers
- Transferring Learning Trajectories of Neural NetworksDaiki ChijiwaICLR 2024 · 4 citations
- TripLe: Revisiting Pretrained Model Reuse and Progressive Learning for Efficient Vision Transformer Scaling and SearchingCheng Fu, Hanxian Huang, Zixuan Jiang, Yun Ni et al.ICCV 2023 · 5 citations
- Quiver: An Informed Storage Cache for Deep LearningAbhishek Vijaya Kumar, Muthian SivathanuFAST 2020 · 91 citations
- Modularizing while Training: A New Paradigm for Modularizing DNN ModelsBinhang Qi, Hailong Sun, Hongyu Zhang, Ruobing Zhao et al.ICSE 2024 · 3 citations
- On Warm-Starting Neural Network TrainingJordan T. Ash, Ryan P. AdamsNeurIPS 2020 · 288 citations
