Cerebro: A Data System for Optimized Deep Learning Model Selection
Supun Nakandala, Yuhao Zhang, Arun Kumar
摘要
Deep neural networks (deep nets) are revolutionizing many machine learning (ML) applications. But there is a major bottleneck to wider adoption: the pain and resource intensiveness of model selection. This empirical process involves exploring deep net architectures and hyper-parameters, often requiring hundreds of trials. Alas, most ML systems focus on training one model at a time, reducing throughput and raising overall resource costs; some also sacrifice reproducibility. We present Cerebro, a new data system to raise deep net model selection throughput at scale without raising resource costs and without sacrificing reproducibility or accuracy. Cerebro uses a new parallel SGD execution strategy we call model hopper parallelism that hybridizes task-and data-parallelism to mitigate the cons of these prior paradigms and offer the best of both worlds. Experiments on large ML benchmark datasets show that Cerebro offers 3x to 10x runtime savings relative to data-parallel systems like Horovod and Parameter Server and up to 8x memory/storage savings or up to 100x network savings relative to task-parallel systems. Cerebro also supports heterogeneous resources and fault tolerance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 被引用 142 次
- Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic ParallelismXupeng Miao, Yujie Wang, Youhe Jiang, Chunan Shi 等VLDB 2023 · 被引用 113 次
- VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space DecompositionYang Li, Yu Shen, Wentao Zhang, Jiawei Jiang 等VLDB 2021 · 被引用 55 次
- SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model DebuggingSvetlana Sagadeeva, Matthias BoehmSIGMOD 2021 · 被引用 45 次
- COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy CompressionSian Jin, Chengming Zhang, Xintong Jiang, Yunhe Feng 等VLDB 2022 · 被引用 39 次
它引用的顶会 Paper1
相关 Paper
- Distributed Deep Learning on Data Systems: A Comparative Analysis of ApproachesYuhao Zhang, Frank Mcquillan, Nandish Jayaram, Nikhil Kak 等VLDB 2021 · 被引用 35 次
- Hippo: Sharing Computations in Hyper-Parameter OptimizationAhnjae Shin, Joo Seong Jeong, Do Yoon Kim, Soyoung Jung 等VLDB 2022 · 被引用 6 次
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu 等OSDI 2023 · 被引用 211 次
- TensorSocket: Shared Data Loading for Deep Learning TrainingTies Robroek, Neil Kim Nielsen, Pinar TözünSIGMOD 2026 · 被引用 4 次
- Herring: rethinking the parameter server at scale for the cloudIndu Thangakrishnan, Derya Cavdar, Can Karakus, Piyush Ghai 等SC 2020 · 被引用 11 次
