TrainBox: An Extreme-Scale Neural Network Training Server Architecture by Systematically Balancing Operations
Pyeongsu Park, Heetaek Jeong, Jangwoo Kim
Abstract
Neural network is a major driving force of another golden age of computing; the computer architects have proposed specialized accelerators (e.g., TPU), high-speed interconnects (e.g., NVLink), and algorithms (e.g., ring-based reduction) to efficiently support neural network applications. As a result, they achieve orders of magnitude higher efficiency for neural network computation and inter-accelerator communication over traditional computing platforms.In this paper, we identify that the emerging platforms have shifted the performance bottleneck of neural network from model computation and inter-accelerator communication to data preparation. Although overlapping data preparation and the others has hidden the preparation overhead, the higher input processing demands of emerging platforms start to reverse the situation; at scale, data preparation requires an infeasible amount of the host-side CPU, memory, and PCIe resources. Our detailed analysis reveals that this heavy resource consumption comes from data transformation for neural network specific formats, and buffering for communication among devices.Therefore, we propose a scalable neural network server architecture by balancing data preparation and the others. To achieve extreme scalability, our design relies on a scalable device array, rather than the limited host resources, with three key ideas. First, we offload CPU-intensive operations to the customized data preparation accelerators to scale the training performance regardless of the host-side CPU performance. Second, we apply direct inter-device communication to eliminate unnecessary data copies and reduce the pressure on the host memory. Lastly, we cluster underlying devices considering unique communication patterns of the neural network processing and interconnect characteristics to efficiently utilize aggregated interconnect band-width. Our evaluation shows that the proposed architecture achieves 44.4× higher training throughput on average over a naively extended server architecture with 256 neural network accelerators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2fc4cd32-2af1-4036-8676-e79e53dcaa1dCited by top-tier papers8
- FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data PipelineTaegeon Um, Byungsoo Oh, Byeongchan Seo, Minhyeok Kweun et al.VLDB 2023 · 45 citations
- Refurbish Your Training Data: Reusing Partially Augmented Samples for Faster Deep Neural Network TrainingGyewon Lee, Irene Lee, Hyeonmin Ha, Kyung-Geun Lee et al.USENIX ATC 2021 · 25 citations
- Pecan: Cost-Efficient ML Data Preprocessing with Automatic Transformation Ordering and Hybrid PlacementDan Graur, Oto Mraz, Muyu Li, Mohammad Sepehr Pourghannad et al.USENIX ATC 2024 · 18 citations
- RAP: Resource-aware Automated GPU Sharing for Multi-GPU Recommendation Model Training and Input PreprocessingZheng Wang, Yuke Wang, Jiaqi Deng, Da Zheng et al.ASPLOS 2024 · 9 citations
- PreSto: An In-Storage Data Preprocessing System for Training Recommendation ModelsYunjae Lee, Hyeseong Kim, Minsoo RhuISCA 2024 · 8 citations
Related papers
- Beyond Inference: Performance Analysis of DNN Server Overheads for Computer VisionAhmed F. AbouElhamayed, Susanne Balle, Deshanand P. Singh, Mohamed S. AbdelfattahDAC 2024 · 3 citations
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 67 citations
- SiP-ML: high-bandwidth optical network interconnects for machine learning trainingMehrdad Khani Shirkoohi, Manya Ghobadi, Mohammad Alizadeh, Ziyi Zhu et al.SIGCOMM 2021 · 94 citations
- Lynx: A SmartNIC-driven Accelerator-centric Architecture for Network ServersMaroun Tork, Lina Maudlej, Mark SilbersteinASPLOS 2020 · 64 citations
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
