Cerebro: A Data System for Optimized Deep Learning Model Selection
Supun Nakandala, Yuhao Zhang, Arun Kumar
Abstract
Deep neural networks (deep nets) are revolutionizing many machine learning (ML) applications. But there is a major bottleneck to wider adoption: the pain and resource intensiveness of model selection. This empirical process involves exploring deep net architectures and hyper-parameters, often requiring hundreds of trials. Alas, most ML systems focus on training one model at a time, reducing throughput and raising overall resource costs; some also sacrifice reproducibility. We present Cerebro, a new data system to raise deep net model selection throughput at scale without raising resource costs and without sacrificing reproducibility or accuracy. Cerebro uses a new parallel SGD execution strategy we call model hopper parallelism that hybridizes task-and data-parallelism to mitigate the cons of these prior paradigms and offer the best of both worlds. Experiments on large ML benchmark datasets show that Cerebro offers 3x to 10x runtime savings relative to data-parallel systems like Horovod and Parameter Server and up to 8x memory/storage savings or up to 100x network savings relative to task-parallel systems. Cerebro also supports heterogeneous resources and fault tolerance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03e1e87b-56ab-4b7c-8574-b6653e1b6a43Cited by top-tier papers21
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 142 citations
- Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic ParallelismXupeng Miao, Yujie Wang, Youhe Jiang, Chunan Shi et al.VLDB 2023 · 113 citations
- VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space DecompositionYang Li, Yu Shen, Wentao Zhang, Jiawei Jiang et al.VLDB 2021 · 55 citations
- SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model DebuggingSvetlana Sagadeeva, Matthias BoehmSIGMOD 2021 · 45 citations
- COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy CompressionSian Jin, Chengming Zhang, Xintong Jiang, Yunhe Feng et al.VLDB 2022 · 39 citations
Builds on1
Related papers
- Distributed Deep Learning on Data Systems: A Comparative Analysis of ApproachesYuhao Zhang, Frank Mcquillan, Nandish Jayaram, Nikhil Kak et al.VLDB 2021 · 35 citations
- Hippo: Sharing Computations in Hyper-Parameter OptimizationAhnjae Shin, Joo Seong Jeong, Do Yoon Kim, Soyoung Jung et al.VLDB 2022 · 6 citations
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu et al.OSDI 2023 · 211 citations
- TensorSocket: Shared Data Loading for Deep Learning TrainingTies Robroek, Neil Kim Nielsen, Pinar TözünSIGMOD 2026 · 4 citations
- Herring: rethinking the parameter server at scale for the cloudIndu Thangakrishnan, Derya Cavdar, Can Karakus, Piyush Ghai et al.SC 2020 · 11 citations
