Efficient GPU Memory Management for Nonlinear DNNs
Donglin Yang, Dazhao Cheng
Abstract
Deep neural networks (DNNs) have been widely applied in the field of artificial intelligence, e.g., natural language processing, computer vision, etc. Researchers and industry practitioners typically use GPU to train complex hundred-layers deep networks. However, as the networks going wider and deeper, the limited GPU memory becomes a significant bottleneck, restricting the size of networks to be trained. In the training of DNNs, the intermediate layer outputs are the major contributors to the memory footprint. Offloading and prefetching feature maps is one of the crucial techniques to overcome the GPU memory shortage by utilizing the CPU DRAM as an external buffer for the GPU. However, we find that the layer-by-layer asynchronous approach cannot be effectively applied to the overlap between communication and computation, particularly for nonlinear networks. Furthermore, the default memory management policy could cause high GPU memory fragmentation for the networks with complex nonlinearities. Based on these observations, we adopt an efficient graph analysis and exploit the layered dependency structures to improve the overlap ratio. To achieve minimal memory fragmentation, we design a Group Tensors By Mobility (GTBM) placement policy to allocate tensors on the proposed unified memory pool for data structures with varied data sizes and dynamic dependencies. We implement and evaluate our system, Dymem, on several linear and nonlinear networks. Compared with vDNN and SuperNeurons, our proposed approach can achieve memory cost reduction by up to 31%. The dependency-aware strategy can improve the end-to-end throughput for nonlinear networks by up to 42%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4787e9cc-c464-4932-b685-bd1b51e8ffcdCited by top-tier papers3
- Demand Layering for Real-Time DNN Inference with Minimized Memory UsageMingoo Ji, Saehanseul Yi, Changjin Koo, Sol Ahn et al.RTSS 2022 · 21 citations
- Occamy: Memory-efficient GPU Compiler for DNN InferenceJaeho Lee, Shinnung Jeong, Seungbin Song, Kunwoo Kim et al.DAC 2023 · 3 citations
- Blaze: Holistic Caching for Iterative Data ProcessingWon Wook Song, Jeongyoon Eo, Taegeon Um, Myeongjae Jeon et al.EuroSys 2024 · 1 citation
Related papers
- TSPLIT: Fine-grained GPU Memory Management for Efficient DNN Training via Tensor SplittingXiaonan Nie, Xupeng Miao, Zhi Yang, Bin CuiICDE 2022 · 26 citations
- DeepUM: Tensor Migration and Prefetching in Unified MemoryJaehoon Jung, Jinpyo Kim, Jaejin LeeASPLOS 2023 · 33 citations
- Enabling Large Dynamic Neural Network Training with Learning-based Memory ManagementJie Ren, Dong Xu, Shuangyan Yang, Jiacheng Zhao et al.HPCA 2024 · 11 citations
- Coop: Memory is not a CommodityJianhao Zhang, Shihan Ma, Peihong Liu, Jinhui YuanNeurIPS 2023 · 11 citations
- Optimal Gradient Checkpoint Search for Arbitrary Computation GraphsJianwei Feng, Dong HuangCVPR 2021
