TSPLIT: Fine-grained GPU Memory Management for Efficient DNN Training via Tensor Splitting
Xiaonan Nie, Xupeng Miao, Zhi Yang, Bin Cui
摘要
Since Deep Neural Networks (DNNs) are deeper and larger, performing DNNs training on existing accelerators (e.g., GPUs) is challenging due to their limited device memory capacity. Existing memory management systems reduce the mem-ory footprint via tensor offloading and recomputing. However, this coarse-grained, one-tensor-at-a-time memory management often incurs high peak GPU memory usage and cannot fully utilize available hardware resources (e.g., PCIe). In this paper, we propose TSPLIT, a fine-grained DNN memory management system that breaks apart memory bottlenecks while maintaining the efficiency of DNNs training. TSPLIT achieves this by proposing a model-guided approach to holistically exploit the tensor-split and its joint optimization with out-of-core execution methods (via offload and recompute). We further provide an efficient implementation of TSPLIT with proposed splittable tensor abstraction, profiling-based planner, and optimized DNN runtime. Evaluations on 6 DNN models show that compared to vDNN and SuperNeurons, TSPLIT can achieve maximum model scale up to 10.5× and 3.1 x and throughput improved up to 4.7× and 2.7 × under the same memory over-subscription, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic ParallelismXupeng Miao, Yujie Wang, Youhe Jiang, Chunan Shi 等VLDB 2023 · 被引用 113 次
- FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device PlacementXiaonan Nie, Xupeng Miao, Zilong Wang, Zichao Yang 等SIGMOD 2023 · 被引用 40 次
- PQCache: Product Quantization-based KVCache for Long Context LLM InferenceHailin Zhang, Xiaodong Ji, Yilin Chen, Fangcheng Fu 等SIGMOD 2025 · 被引用 13 次
- Ratel: Optimizing Holistic Data Movement to Fine-tune 100B Model on a Consumer GPUChangyue Liao, Mo Sun, Zihan Yang, Jun Xie 等ICDE 2025 · 被引用 4 次
- Occamy: Memory-efficient GPU Compiler for DNN InferenceJaeho Lee, Shinnung Jeong, Seungbin Song, Kunwoo Kim 等DAC 2023 · 被引用 3 次
相关 Paper
- Tensor Movement Orchestration in Multi-GPU Training SystemsShao-Fu Lin, Yi-Jung Chen, Hsiang-Yun Cheng, Chia-Lin YangHPCA 2023 · 被引用 5 次
- Efficient GPU Memory Management for Nonlinear DNNsDonglin Yang, Dazhao ChengHPDC 2020 · 被引用 18 次
- Zico: Efficient GPU Memory Sharing for Concurrent DNN TrainingGangmuk Lim, Jeongseob Ahn, Wencong Xiao, Youngjin Kwon 等USENIX ATC 2021 · 被引用 65 次
- GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory StitchingCong Guo, Rui Zhang, Jiale Xu, Jingwen Leng 等ASPLOS 2024 · 被引用 30 次
- T-Control: An Efficient Dynamic Tensor Rematerialization System for DNN TrainingZehua Wang, Junmin Xiao, Xiaochuan Deng, Huibing Wang 等ASPLOS 2026 · 被引用 2 次
