HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training
Guicheng Qi, Junwei Su, Liqi Yang, Tao Li, Tingwen Xie, Yerui Sun, Yuchen Xie, Chuan Wu
Abstract
As large neural network models (e.g., LLMs) grow in scale, single-cluster resources become insufficient, making cross-cluster distributed training essential. Cross-cluster training is challenging: hardware heterogeneity complicates load balancing and parallelization strategy and introduces hardware compatibility issues in implementation; cross-cluster communication bottlenecks severely impact training throughput. We present HetAuto, an automatic parallelization system for efficient cross-cluster heterogeneous large model training. HetAuto contributes three key innovations: (1) a principle-guided MCTS algorithm with a random forest-enhanced cost model that efficiently searches parallelization strategies and quickly evaluates their performance under heterogeneous configurations; (2) cross-cluster communication optimizations including Virtual-1F1B scheduling that overlaps communication with computation and an optimized resharding strategy for inter-stage communication; and (3) a unified API enabling seamless integration of diverse accelerators. We evaluate HetAuto across 4 different clusters with up to 736 heterogeneous devices. The evaluation results show that HetAuto achieves up to 1.57× training throughput improvement over representative baselines, and strikes an efficient balance between solution quality and search overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73e3de93-a47e-4064-9286-18f23f313971Builds on24
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- DAPPLE: a pipelined data parallel approach for training large modelsShiqing Fan, Yi Rong, Chen Meng, Zongyan Cao et al.PPoPP 2021 · 224 citations
Related papers
- AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM TrainingYuanyuan Wang, Nana Tang, Yuyang Wang, Shu Pan et al.HPCA 2026
- HeteCCL: Synthesizing Near-Optimal Collective Communication Schedules for Heterogeneous GPU ClustersChenyang Hei, Jiayi Li, Jiamin Cao, Chengxi Gao et al.NSDI 2026 · 4 citations
- HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program SynthesisShiwei Zhang, Lansong Diao, Chuan Wu, Zongyan Cao et al.EuroSys 2024 · 16 citations
- Accelerating Multi-modal LLM Training with Adaptive Model Placement and ParallelizationYiming Yin, Shaohuai Shi, Qiang Wang, Xiaowen ChuINFOCOM 2026
- HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU ClustersAntian Liang, Zhigang Zhao, Kai Zhang, Xuri Shi et al.EuroSys 2026 · 1 citation
