Enabling Efficient Large-Scale Deep Learning Training with Cache Coherent Disaggregated Memory Systems
Zixuan Wang, Joonseop Sim, Euicheol Lim, Jishen Zhao
Abstract
Modern deep learning (DL) training is memory-consuming, constrained by the memory capacity of each computation component and cross-device communication bandwidth. In response to such constraints, current approaches include increasing parallelism in distributed training and optimizing inter-device communication. However, model parameter communication is becoming a key performance bottleneck in distributed DL training. To improve parameter communication performance, we propose COARSE, a disaggregated memory extension for distributed DL training. COARSE is built on modern cache-coherent interconnect (CCI) protocols and MPI-like collective communication for synchronization, to allow low-latency and parallel access to training data and model parameters shared among worker GPUs. To enable high bandwidth transfers between GPUs and the disaggregated memory system, we propose a decentralized parameter communication scheme to decouple and localize parameter synchronization traffic. Furthermore, we propose dynamic tensor routing and partitioning to fully utilize the non-uniform serial bus bandwidth varied across different cloud computing systems. Finally, we design a deadlock avoidance and dual synchronization to ensure high-performance parameter synchronization. Our evaluation shows that COARSE achieves up to 48.3% faster DL training compared to the state-of-the-art MPI AllReduce communication.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- Low-Overhead General-Purpose Near-Data Processing in CXL Memory ExpandersHyungkyu Ham, Jeongmin Hong, Geonwoo Park, Yunseon Shin et al.MICRO 2024 · 26 citations
- Scaling Up Memory Disaggregated Applications with SMARTFeng Ren, Mingxing Zhang, Kang Chen, Huaxia Xia et al.ASPLOS 2024 · 16 citations
- EdgeMove: Pipelining Device-Edge Model Training for Mobile IntelligenceZeqian Dong, Qiang He, Feifei Chen, Hai Jin et al.WWW 2023 · 12 citations
- CAESAR: Coherence-Aided Elective and Seamless Alternative Routing via on-chip FPGAShahin Roozkhosh, Denis Hoornaert, Renato MancusoRTSS 2022 · 10 citations
Related papers
- Efficient Tensor Offloading for Large Deep-Learning Model Training based on Compute Express LinkDong Xu, Yuan Feng, Kwangsik Shin, Daewoo Kim et al.SC 2024 · 16 citations
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 67 citations
- Mobius: Fine Tuning Large-Scale Models on Commodity GPU ServersYangyang Feng, Minhui Xie, Zijie Tian, Shuo Wang et al.ASPLOS 2023 · 29 citations
- SkipReduce: (Interconnection) Network Sparsity to Accelerate Distributed Machine LearningHans Kasan, Dennis Abts, Jungwook Choi, John KimMICRO 2025 · 1 citation
- Near-Optimal Topology-adaptive Parameter Synchronization in Distributed DNN TrainingZhe Zhang, Chuan Wu, Zongpeng LiINFOCOM 2021 · 14 citations
