Metis: Fast Automatic Distributed Training on Heterogeneous GPUs
Taegeon Um, Byungsoo Oh, Minyoung Kang, Woo-Yeon Lee, Goeun Kim, Dongseob Kim, Youngtaek Kim, Mohd Muzzammil, Myeongjae Jeon
Abstract
As deep learning model sizes expand and new GPUs are released every year, the need for distributed training on heterogeneous GPUs rises to fully harness under-utilized low-end GPUs and reduce the cost of purchasing expensive high-end GPUs. In this paper, we introduce Metis, a system designed to automatically find efficient parallelism plans for distributed training on heterogeneous GPUs. Metis holistically optimizes several key system components, such as profiler, cost estimator, and planner, which were limited to single GPU types, to now efficiently leverage compute powers and memory capacities of diverse GPU types. This enables Metis to achieve fine-grained distribution of training workloads across heterogeneous GPUs, improving resource efficiency. However, the search space designed for automatic parallelism in this complexity would be prohibitively expensive to navigate.
To address this issue, Metis develops a new search algorithm that efficiently prunes large search spaces and balances loads with heterogeneity-awareness, while preferring data parallelism over tensor parallelism within a pipeline stage to take advantage of its superior computation and communication trade-offs. Our evaluation with three large models (GPT-3, MoE, and Wide-Resnet) on combinations of three types of GPUs demonstrates that Metis finds better parallelism plans than traditional methods with 1.05 ∼ 8.43× training speed-up, while requiring less profiling searching time. Compared to the oracle planning that delivers the fastest parallel training, Metis finds near-optimal solutions while reducing profiling and search overheads by orders of magnitude.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3b3ec1e-775a-42e4-8425-83026f8de01dCited by top-tier papers15
- GREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at ScaleTianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang et al.USENIX ATC 2025 · 19 citations
- SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingJiamin Cao, Shangfeng Shi, Jiaqi Gao, Weisen Liu et al.SIGCOMM 2025 · 15 citations
- Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data AnnotationsHaoyang Li, Fangcheng Fu, Hao Ge, Sheng Lin et al.OSDI 2026 · 7 citations
- FlexPipe: Maximizing Training Efficiency for Transformer-based Models with Variable-Length InputsHairui Zhao, Qi Tian, Hongliang Li, Zizhong ChenUSENIX ATC 2025 · 6 citations
- Mario: Near Zero-cost Activation Checkpointing in Pipeline ParallelismWeijian Liu, Mingzhen Li, Guangming Tan, Weile JiaPPoPP 2025 · 4 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee et al.OSDI 2020 · 286 citations
Related papers
- Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic ParallelismZizhao Mo, Jianxiong Liao, Huanle Xu, Zhi Zhou et al.SC 2025 · 3 citations
- Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU ClustersWenZheng Zhang, Yang Hu, Jing Shi, Xiaoying BaiAAAI 2025 · 5 citations
- Mobius: Fine Tuning Large-Scale Models on Commodity GPU ServersYangyang Feng, Minhui Xie, Zijie Tian, Shuo Wang et al.ASPLOS 2023 · 29 citations
- Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningLianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang et al.OSDI 2022 · 75 citations
- FASOP: Fast yet Accurate Automated Search for Optimal Parallelization of Transformers on Heterogeneous GPU ClustersSunyeol Hwang, Eungyeong Lee, Hongseok Oh, Youngmin YiHPDC 2024 · 4 citations
