Sentinel: Efficient Tensor Migration and Allocation on Heterogeneous Memory Systems for Deep Learning
Jie Ren, Jiaolin Luo, Kai Wu, Minjia Zhang, Hyeran Jeon, Dong Li
Abstract
Memory capacity is a major bottleneck for training deep neural networks (DNN). Heterogeneous memory (HM) combining fast and slow memories provides a promising direction to increase memory capacity. However, HM imposes challenges on tensor migration and allocation for high performance DNN training. Prior work heavily relies on DNN domain knowledge, unnecessarily causes tensor migration due to page-level false sharing, and wastes fast memory space. We present Sentinel, a software runtime system that automatically optimizes tensor management on HM. Sentinel uses dynamic profiling, and coordinates operating system (OS) and runtime-level profiling to bridge the semantic gap between OS and applications, which enables tensor-level profiling. This profiling enables co-allocating tensors with similar lifetime and memory access frequency into the same pages. Such fine-grained profiling and tensor collocation avoids unnecessary data movement, improves tensor movement efficiency, and enables larger batch training because of saving in fast memory space. Sentinel reduces fast memory consumption by 80% while retaining comparable performance to fast memory-only system; Sentinel consistently outperforms a state-of-the-art solution on CPU by 37% and two state-of-the-art solutions on GPU by 2x and 21% respectively in training throughput.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a2c38005-a854-444f-bf8c-7bab38c441faCited by top-tier papers26
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase et al.USENIX ATC 2021 · 657 citations
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith et al.SC 2021 · 254 citations
- MEMTIS: Efficient Memory Tiering with Dynamic Page Classification and Page Size DeterminationTaehyung Lee, Sumit Kumar Monga, Changwoo Min, Young Ik EomSOSP 2023 · 67 citations
- Full Parameter Fine-tuning for Large Language Models with Limited ResourcesKai Lv, Yuqing Yang, Tengxiao Liu, Qipeng Guo et al.ACL 2024 · 61 citations
- SmartSAGE: training large-scale graph neural networks using in-storage processing architecturesYunjae Lee, Jinha Chung, Minsoo RhuISCA 2022 · 57 citations
Related papers
- TSPLIT: Fine-grained GPU Memory Management for Efficient DNN Training via Tensor SplittingXiaonan Nie, Xupeng Miao, Zhi Yang, Bin CuiICDE 2022 · 26 citations
- HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program SynthesisShiwei Zhang, Lansong Diao, Chuan Wu, Zongyan Cao et al.EuroSys 2024 · 16 citations
- Orion: Interference-aware, Fine-grained GPU Sharing for ML ApplicationsFoteini Strati, Xianzhe Ma, Ana KlimovicEuroSys 2024 · 96 citations
- Tensor Movement Orchestration in Multi-GPU Training SystemsShao-Fu Lin, Yi-Jung Chen, Hsiang-Yun Cheng, Chia-Lin YangHPCA 2023 · 5 citations
- Merchandiser: Data Placement on Heterogeneous Memory for Task-Parallel HPC Applications with Load-Balance AwarenessZhen Xie, Jie Liu, Jiajia Li, Dong LiPPoPP 2023 · 18 citations
