Lune

ICDE2025Top-tier venue

Ratel: Optimizing Holistic Data Movement to Fine-tune 100B Model on a Consumer GPU

Changyue Liao, Mo Sun, Zihan Yang, Jun Xie, Kaiqi Chen, Binhang Yuan, Fei Wu, Zeke Wang

2025Year
4Citations
1Top-tier citations

Abstract

Nowadays, AI researchers become more and more interested in fine-tuning a pre-trained LLM, whose size has grown to up to over 100B parameters, for their downstream tasks. One approach to fine-tune such huge models is to aggregate device memory from many GPUs. However, this approach introduces prohibitive costs for most data scientists with a limited budget for high-end GPU servers. In this paper, we focus on LLM fine-tuning on a single consumer-grade GPU in a commodity server with limited main memory capacity, which is accessible to most AI researchers. In such a scenario, existing offloadingbased methods fail to fine-tune an LLM efficiently due to a lack of holistic intra-server tensor movement management. To this end, we present Ratel, a low-cost, high-performance deep learning training framework that enables efficient 100B-scale model fine-tuning on a commodity server with a consumergrade GPU and limited main memory capacity. The key idea is to add holistic offloading traffic as an optimization dimension for 1) active gradient offloading, and 2) holistic traffic-aware activation swapping mechanism. The experimental results show that 1) Ratel is the first to fine-tune a 175B model on an RTX 4090 and 256 GB main memory, 2) Ratel achieves 2.32× throughput than the state-of-the-art baselines when fine-tuning a small 13B model, and 3) Ratel enables a cheap low-end consumer GPU to have higher cost-effectiveness than a DGX-A100 cluster when fine-tuning a 175B model.

  • Equal contribution. 1 Size of an LLM is defined by the number of parameters. We use "100B model" to represent a model with 100 billion parameters in this paper.

In this paper, we aim to explore whether it is feasible to efficiently fine-tune a 100B-scale LLM on a single consumergrade 4090 GPU ($1600, up to 24 GB device memory) with limited main memory capacity (256 GB). Such a solution would be attractive to researchers who seek to minimize LLM fine-tuning costs. To do so, the existing low-cost works [31]-[36] offload the tensors during the fine-tuning process from GPU memory to NVMe memory to maximize the trainable model size. However, we identify that these SSD-equipped systems suffer from two severe issues: low throughput and small maximum trainable model size.

• Offloading Activation Tensors to SSDs. LLM training consists of two types of tensors, namely activations and model states. Existing systems like FlashNeuron [37] only offload activations to SSDs and keep model states in GPU memory. We find that keeping model states in GPU memory severely limits the trainable model size. For example, Flash-Neuron can only fine-tune a 1.55B model on RTX 4090, while fine-tuning a 175B model (a typical size of 100Bscale models [2], [38]) requires 2.45 TB of GPU memory, which far exceeds the memory capacity of GPUs. • Offloading Model State Tensors to SSDs. Existing systems like ZeRO-Infinity [39] and Colossal-AI [40] offload model states to NVMe SSDs [41]- [43] to enlarge the trainable model size. We identify that these systems have three issues in model fine-tuning, as shown in Figure 1a. First, these systems suffer from low GPU utilization, mainly because they execute the synchronous out-of-core CPU Adam 2 [44] in the optimizer stage where the GPU is idle. This stage takes 30% 60% of a training iteration. Second, these systems only offload the inter-transformer block activations (6% of total activations) to main memory and recompute the rest of the activations. Such an offloading method leads to 5.7 seconds (22% of the backward stage) of additional GPU recomputation overhead in the backward stage where PCIe bandwidth is underutilized. Third, these systems only offload activations to main memory, thus requiring a large amount of main memory to fine-tune an LLM. We estimate that ZeRO-Infinity requires 1.1 TB main memory to fine-tune the 175B model, while most commodity servers equip only 128 GB 1 TB main memory.

2 Here we refer to the out-of-core optimizer as optimizer executing on CPU instead of GPU (i.e., "in-core optimizer"). There are also works such as Angel-PTM [31] that present an asynchronous out-of-core optimizer. However, the asynchronous optimizer updating policy might affect model training convergence. Therefore, they are beyond the scope of this paper.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers1

Ask how each one uses it

Builds on49

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines