Lune

ICDE2025顶会

Ratel: Optimizing Holistic Data Movement to Fine-tune 100B Model on a Consumer GPU

Changyue Liao, Mo Sun, Zihan Yang, Jun Xie, Kaiqi Chen, Binhang Yuan, Fei Wu, Zeke Wang

2025年份
4被引次数
1顶会引用

摘要

Nowadays, AI researchers become more and more interested in fine-tuning a pre-trained LLM, whose size has grown to up to over 100B parameters, for their downstream tasks. One approach to fine-tune such huge models is to aggregate device memory from many GPUs. However, this approach introduces prohibitive costs for most data scientists with a limited budget for high-end GPU servers. In this paper, we focus on LLM fine-tuning on a single consumer-grade GPU in a commodity server with limited main memory capacity, which is accessible to most AI researchers. In such a scenario, existing offloadingbased methods fail to fine-tune an LLM efficiently due to a lack of holistic intra-server tensor movement management. To this end, we present Ratel, a low-cost, high-performance deep learning training framework that enables efficient 100B-scale model fine-tuning on a commodity server with a consumergrade GPU and limited main memory capacity. The key idea is to add holistic offloading traffic as an optimization dimension for 1) active gradient offloading, and 2) holistic traffic-aware activation swapping mechanism. The experimental results show that 1) Ratel is the first to fine-tune a 175B model on an RTX 4090 and 256 GB main memory, 2) Ratel achieves 2.32× throughput than the state-of-the-art baselines when fine-tuning a small 13B model, and 3) Ratel enables a cheap low-end consumer GPU to have higher cost-effectiveness than a DGX-A100 cluster when fine-tuning a 175B model.

  • Equal contribution. 1 Size of an LLM is defined by the number of parameters. We use "100B model" to represent a model with 100 billion parameters in this paper.

In this paper, we aim to explore whether it is feasible to efficiently fine-tune a 100B-scale LLM on a single consumergrade 4090 GPU ($1600, up to 24 GB device memory) with limited main memory capacity (256 GB). Such a solution would be attractive to researchers who seek to minimize LLM fine-tuning costs. To do so, the existing low-cost works [31]-[36] offload the tensors during the fine-tuning process from GPU memory to NVMe memory to maximize the trainable model size. However, we identify that these SSD-equipped systems suffer from two severe issues: low throughput and small maximum trainable model size.

• Offloading Activation Tensors to SSDs. LLM training consists of two types of tensors, namely activations and model states. Existing systems like FlashNeuron [37] only offload activations to SSDs and keep model states in GPU memory. We find that keeping model states in GPU memory severely limits the trainable model size. For example, Flash-Neuron can only fine-tune a 1.55B model on RTX 4090, while fine-tuning a 175B model (a typical size of 100Bscale models [2], [38]) requires 2.45 TB of GPU memory, which far exceeds the memory capacity of GPUs. • Offloading Model State Tensors to SSDs. Existing systems like ZeRO-Infinity [39] and Colossal-AI [40] offload model states to NVMe SSDs [41]- [43] to enlarge the trainable model size. We identify that these systems have three issues in model fine-tuning, as shown in Figure 1a. First, these systems suffer from low GPU utilization, mainly because they execute the synchronous out-of-core CPU Adam 2 [44] in the optimizer stage where the GPU is idle. This stage takes 30% 60% of a training iteration. Second, these systems only offload the inter-transformer block activations (6% of total activations) to main memory and recompute the rest of the activations. Such an offloading method leads to 5.7 seconds (22% of the backward stage) of additional GPU recomputation overhead in the backward stage where PCIe bandwidth is underutilized. Third, these systems only offload activations to main memory, thus requiring a large amount of main memory to fine-tune an LLM. We estimate that ZeRO-Infinity requires 1.1 TB main memory to fine-tune the 175B model, while most commodity servers equip only 128 GB 1 TB main memory.

2 Here we refer to the out-of-core optimizer as optimizer executing on CPU instead of GPU (i.e., "in-core optimizer"). There are also works such as Angel-PTM [31] that present an asynchronous out-of-core optimizer. However, the asynchronous optimizer updating policy might affect model training convergence. Therefore, they are beyond the scope of this paper.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper49

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖