Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
Ziqi Yuan, Haoyang Zhang, Yirui Eric Zhou, Apoorve Mohan, I-Hsin Chung, Seetharami Seelam, Jian Huang
Abstract
We present the design and implementation of a new lifetime-aware tensor offloading framework for GPU memory expansion using low-cost PCIe-based solid-state drives (SSDs). Our framework, TERAIO, is developed explicitly for large language model (LLM) training with multiple GPUs and multiple SSDs. Its design is driven by our observation that the active tensors take only a small fraction (1.7% on average) of allocated GPU memory in each LLM training iteration, the inactive tensors are usually large and will not be used for a long period of time, creating ample opportunities for offloading/prefetching tensors to/from slow SSDs without stalling the GPU training process. TERAIO accurately estimates the lifetime (active period of time in GPU memory) of each tensor with the profiling of the first few iterations in the training process. With the tensor lifetime analysis, TERAIO will generate an optimized tensor offloading/prefetching plan and integrate it into the compiled LLM program via PyTorch. TERAIO has a runtime tensor migration engine to execute the offloading/prefetching plan via GPUDirect storage, which allows direct tensor migration between GPUs and SSDs for alleviating the CPU bottleneck and maximizing the SSD bandwidth utilization. In comparison with state-of-the-art studies such as ZeRO-Offload and ZeRO-Infinity, we show that TERAIO improves the training performance of various LLMs by 1.47× on average, and achieves 80.7% of the ideal performance assuming unlimited GPU memory. tensors (GPU memory, host memory, or SSDs), TERAIO indexes tensors with their identification numbers using hash maps.
We implement the core components of TERAIO based on PyTorch. Therefore, TERAIO does not require any code modifications to LLM training programs. To evaluate the efficiency of TERAIO, we train a set of Llama and Granite models with different batch sizes and sequence lengths using TorchTitan [22] on a GPU server that has two NVIDIA H100 GPUs and eight PCIe-based SSDs. In comparison with state-of-the-art offloading solutions and , TERAIO improves the training performance by 1.47× on average, achieves 80.7% of the ideal performance assuming unlimited GPU on-board memory, and delivers 1.45× improvement on cost efficiency for LLM training. In summary, we make the following contributions.
• We conduct a quantitative characterization study of tensor memory usage when training different LLMs on multiple GPUs, and show that the high compute intensity of modern LLMs provide rich opportunities for tensor offloading.
• We develop a lightweight tensor lifetime profiler based on PyTorch, which can learn tensor activity patterns for multi-GPU LLM training.
• We design a lifetime-aware tensor migration planning algorithm that optimizes offloading/prefetching decisions based on tensor activity patterns, GPU memory capacity, and the available migration bandwidth.
• We implement a transparent tensor migration engine that enables direct data transfer between GPU and SSDs, alleviating the scalability bottleneck on the host.
• We conduct a thorough evaluation of TERAIO with the training of various LLMs, demonstrating significant improvement on training performance and cost efficiency, compared to state-of-the-art offloading solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 332d2514-e8db-4d8b-bb28-b88447d409bdBuilds on19
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase et al.USENIX ATC 2021 · 657 citations
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith et al.SC 2021 · 254 citations
- NetLLM: Adapting Large Language Models for NetworkingDuo Wu, Xianda Wang, Yaqi Qiao, Zhi Wang et al.SIGCOMM 2024 · 162 citations
Related papers
- SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model TrainingKun Wu, Jeongmin Brian Park, Xiaofan Zhang, Mert Hidayetoglu et al.DAC 2025 · 3 citations
- STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal PlanningZixiao Huang, Junhao Hu, Hao Lin, Chunyang Zhu et al.EuroSys 2026 · 1 citation
- TurboBus: Pooling PCIe Bandwidth for LLM Workloads via Scale-Up FabricsXinyu Yang, Kaiqiang Xu, Kai ChenSIGCOMM 2026
- Practical Offloading for Fine-Tuning LLM on Commodity GPU via Learned Sparse ProjectorsSiyuan Chen, Zhuofeng Wang, Zelong Guan, Yudong Liu et al.AAAI 2025 · 3 citations
- Enabling Large Dynamic Neural Network Training with Learning-based Memory ManagementJie Ren, Dong Xu, Shuangyan Yang, Jiacheng Zhao et al.HPCA 2024 · 11 citations
