An Oracle for Guiding Large-Scale Model/Hybrid Parallel Training of Convolutional Neural Networks
Albert Njoroge Kahira, Truong Thao Nguyen, Leonardo Bautista-Gomez, Ryousei Takano, Rosa M. Badia, Mohamed Wahib
Abstract
Deep Neural Network (DNN) frameworks use distributed training to enable faster time to convergence and alleviate memory capacity limitations when training large models and/or using high dimension inputs. With the steady increase in datasets and model sizes, model/hybrid parallelism is deemed to have an important role in the future of distributed training of DNNs. We analyze the compute, communication, and memory requirements of Convolutional Neural Networks (CNNs) to understand the trade-offs between different parallelism approaches on performance and scalability. We leverage our model-driven analysis to be the basis for an oracle utility which can help in detecting the limitations and bottlenecks of different parallelism approaches at scale. We evaluate the oracle on six parallelization strategies, with four CNN models and multiple datasets (2D and 3D), on up to 1024 GPUs. The results demonstrate that the oracle has an average accuracy of about 86.74% when compared to empirical results, and as high as 97.57% for data parallelism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 477f4063-7ff8-4bd9-a6a5-c1344780f790Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Convolutional neural network training with distributed K-FACJ. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu et al.SC 2020 · 26 citations
- Extracting clean performance models from tainted programsMarcin Copik, Alexandru Calotoiu, Tobias Grosser, Nicolas Wicki et al.PPoPP 2021 · 16 citations
- Scaling distributed deep learning workloads beyond the memory capacity with KARMAMohamed Wahib, Haoyu Zhang, Truong Thao Nguyen, Aleksandr Drozd et al.SC 2020 · 15 citations
Related papers
- Harmony: Overcoming the hurdles of GPU memory capacity to train massive DNN models on commodity serversYoujie Li, Amar Phanishayee, Derek Murray, Jakub Tarnawski et al.VLDB 2022 · 28 citations
- Hypertron: Efficiently Scaling Large Models by Exploring High-Dimensional Parallelization SpaceShigang Li, Jingkun Dong, Jihao Chen, Zhi Ma et al.SC 2025 · 2 citations
- DAPPLE: a pipelined data parallel approach for training large modelsShiqing Fan, Yi Rong, Chen Meng, Zongyan Cao et al.PPoPP 2021 · 224 citations
- Aceso: Efficient Parallel DNN Training through Iterative Bottleneck AlleviationGuodong Liu, Youshan Miao, Zhiqi Lin, Xiaoxiang Shi et al.EuroSys 2024 · 16 citations
- Training Acceleration for Deep Neural Networks: A Hybrid Parallelization StrategyZihao Zeng, Chubo Liu, Zhuo Tang, Wanli Chang et al.DAC 2021 · 13 citations
