Espresso: Cost-Efficient Large Model Training by Exploiting GPU Heterogeneity in the Cloud
Qiannan Zhou, Fei Xu, Lingxuan Weng, Ruixing Li, Xudong Wu, Li Chen, Zhi Zhou, Fangming Liu
摘要
As Transformer-based models deepen and datasets expand, training large models demands numerous accelerators, particularly GPUs, bringing high cloud expenses. However, conventional homogeneous resource provisioning is inefficient due to limited cloud resources and low GPU utilization. This challenge necessitates heterogeneous GPU provisioning for training in clouds. Current research on large model training often focuses on load balancing of stages, neglecting the varying computing and memory demands across stages. Additionally, the allocation of heterogeneous G PU s for training has surprisingly received little attention. This paper introduces Espresso, a cost-efficient GPU provisioning framework that unifies the heterogeneous GPU allocation (GPU allocator) and adequate stage placement (stage placer) for large model training in the cloud. Specifically, the GPU allocator proposes a cost tree-based provisioning strategy to prioritize searching allocation plans with lower costs and reduce unnecessary branches by multi-dimensional pruning methods. The resource-aware stage placer further devises a compute-memory ratio to optimize communication and computation efficiency during training. We have open-sourced a prototype of Espresso and conducted prototype experiments on four representative large models in public clouds. Extensive experiment results demonstrate that Espresso guarantees the performance for large model training while saving costs by up to 49.8 % compared to state-of-the-art solutions, yet with acceptable runtime overhead.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-DesignChunyu Xue, Weihao Cui, Quan Chen, Chen Chen 等EuroSys 2026
- EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUsMingzhen Li, Wencong Xiao, Hailong Yang, Biao Sun 等SC 2023 · 被引用 16 次
- Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUsYouhe Jiang, Fangcheng Fu, Xiaozhe Yao, Guoliang He 等ICML 2025
- Metis: Fast Automatic Distributed Training on Heterogeneous GPUsTaegeon Um, Byungsoo Oh, Minyoung Kang, Woo-Yeon Lee 等USENIX ATC 2024 · 被引用 81 次
- HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU ClustersAntian Liang, Zhigang Zhao, Kai Zhang, Xuri Shi 等EuroSys 2026 · 被引用 1 次
