Sentinel: Efficient Tensor Migration and Allocation on Heterogeneous Memory Systems for Deep Learning
Jie Ren, Jiaolin Luo, Kai Wu, Minjia Zhang, Hyeran Jeon, Dong Li
摘要
Memory capacity is a major bottleneck for training deep neural networks (DNN). Heterogeneous memory (HM) combining fast and slow memories provides a promising direction to increase memory capacity. However, HM imposes challenges on tensor migration and allocation for high performance DNN training. Prior work heavily relies on DNN domain knowledge, unnecessarily causes tensor migration due to page-level false sharing, and wastes fast memory space. We present Sentinel, a software runtime system that automatically optimizes tensor management on HM. Sentinel uses dynamic profiling, and coordinates operating system (OS) and runtime-level profiling to bridge the semantic gap between OS and applications, which enables tensor-level profiling. This profiling enables co-allocating tensors with similar lifetime and memory access frequency into the same pages. Such fine-grained profiling and tensor collocation avoids unnecessary data movement, improves tensor movement efficiency, and enables larger batch training because of saving in fast memory space. Sentinel reduces fast memory consumption by 80% while retaining comparable performance to fast memory-only system; Sentinel consistently outperforms a state-of-the-art solution on CPU by 37% and two state-of-the-art solutions on GPU by 2x and 21% respectively in training throughput.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper26
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase 等USENIX ATC 2021 · 被引用 657 次
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith 等SC 2021 · 被引用 254 次
- MEMTIS: Efficient Memory Tiering with Dynamic Page Classification and Page Size DeterminationTaehyung Lee, Sumit Kumar Monga, Changwoo Min, Young Ik EomSOSP 2023 · 被引用 67 次
- Full Parameter Fine-tuning for Large Language Models with Limited ResourcesKai Lv, Yuqing Yang, Tengxiao Liu, Qipeng Guo 等ACL 2024 · 被引用 61 次
- SmartSAGE: training large-scale graph neural networks using in-storage processing architecturesYunjae Lee, Jinha Chung, Minsoo RhuISCA 2022 · 被引用 57 次
相关 Paper
- TSPLIT: Fine-grained GPU Memory Management for Efficient DNN Training via Tensor SplittingXiaonan Nie, Xupeng Miao, Zhi Yang, Bin CuiICDE 2022 · 被引用 26 次
- HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program SynthesisShiwei Zhang, Lansong Diao, Chuan Wu, Zongyan Cao 等EuroSys 2024 · 被引用 16 次
- Orion: Interference-aware, Fine-grained GPU Sharing for ML ApplicationsFoteini Strati, Xianzhe Ma, Ana KlimovicEuroSys 2024 · 被引用 96 次
- Tensor Movement Orchestration in Multi-GPU Training SystemsShao-Fu Lin, Yi-Jung Chen, Hsiang-Yun Cheng, Chia-Lin YangHPCA 2023 · 被引用 5 次
- Merchandiser: Data Placement on Heterogeneous Memory for Task-Parallel HPC Applications with Load-Balance AwarenessZhen Xie, Jie Liu, Jiajia Li, Dong LiPPoPP 2023 · 被引用 18 次
