Enabling Efficient Large-Scale Deep Learning Training with Cache Coherent Disaggregated Memory Systems
Zixuan Wang, Joonseop Sim, Euicheol Lim, Jishen Zhao
摘要
Modern deep learning (DL) training is memory-consuming, constrained by the memory capacity of each computation component and cross-device communication bandwidth. In response to such constraints, current approaches include increasing parallelism in distributed training and optimizing inter-device communication. However, model parameter communication is becoming a key performance bottleneck in distributed DL training. To improve parameter communication performance, we propose COARSE, a disaggregated memory extension for distributed DL training. COARSE is built on modern cache-coherent interconnect (CCI) protocols and MPI-like collective communication for synchronization, to allow low-latency and parallel access to training data and model parameters shared among worker GPUs. To enable high bandwidth transfers between GPUs and the disaggregated memory system, we propose a decentralized parameter communication scheme to decouple and localize parameter synchronization traffic. Furthermore, we propose dynamic tensor routing and partitioning to fully utilize the non-uniform serial bus bandwidth varied across different cloud computing systems. Finally, we design a deadlock avoidance and dual synchronization to ensure high-performance parameter synchronization. Our evaluation shows that COARSE achieves up to 48.3% faster DL training compared to the state-of-the-art MPI AllReduce communication.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Low-Overhead General-Purpose Near-Data Processing in CXL Memory ExpandersHyungkyu Ham, Jeongmin Hong, Geonwoo Park, Yunseon Shin 等MICRO 2024 · 被引用 26 次
- Scaling Up Memory Disaggregated Applications with SMARTFeng Ren, Mingxing Zhang, Kang Chen, Huaxia Xia 等ASPLOS 2024 · 被引用 16 次
- EdgeMove: Pipelining Device-Edge Model Training for Mobile IntelligenceZeqian Dong, Qiang He, Feifei Chen, Hai Jin 等WWW 2023 · 被引用 12 次
- CAESAR: Coherence-Aided Elective and Seamless Alternative Routing via on-chip FPGAShahin Roozkhosh, Denis Hoornaert, Renato MancusoRTSS 2022 · 被引用 10 次
相关 Paper
- Efficient Tensor Offloading for Large Deep-Learning Model Training based on Compute Express LinkDong Xu, Yuan Feng, Kwangsik Shin, Daewoo Kim 等SC 2024 · 被引用 16 次
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 被引用 67 次
- Mobius: Fine Tuning Large-Scale Models on Commodity GPU ServersYangyang Feng, Minhui Xie, Zijie Tian, Shuo Wang 等ASPLOS 2023 · 被引用 29 次
- SkipReduce: (Interconnection) Network Sparsity to Accelerate Distributed Machine LearningHans Kasan, Dennis Abts, Jungwook Choi, John KimMICRO 2025 · 被引用 1 次
- Near-Optimal Topology-adaptive Parameter Synchronization in Distributed DNN TrainingZhe Zhang, Chuan Wu, Zongpeng LiINFOCOM 2021 · 被引用 14 次
