AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM Training
Yuanyuan Wang, Nana Tang, Yuyang Wang, Shu Pan, Dingding Yu, Zeyue Wang, Mou Sun, Kejie Fu, Fangyu Wang, Yunchuan Chen, Ning Sun, Fei Yang
Abstract
Heterogeneous clusters with diverse devices mitigate computational and memory burdens in large language model (LLM) training, yet their inherent resource heterogeneity, characterized by divergent computation, memory, and bandwidth capabilities, renders manual parallelization strategy optimization both challenging and time-intensive. Automatic parallelization is critical for scaling complex workloads across heterogeneous architectures. However, previous methodologies face significant inefficiencies. First, insufficient pruning of the parameter initialization space results in impractically large search spaces. Second, the prevailing automatic parallel search strategies exhibit suboptimal performance in load balancing and resource constraint adaptation. Third, dynamic parallel strategy tuning incurs substantial overhead due to redundant latency calculations for operators with unchanged configurations, leading to unnecessary computational costs. Therefore, insufficient search space pruning, suboptimal load/resource adaptation, and redundant latency computation are identified as the major bottlenecks in our research. To address these challenges, we propose AutoHAAP (Automated Heterogeneity-Aware Asymmetric Partitioning), a novel framework incorporating three core innovations: (1) memory-aware initialization to drastically reduce viable search spaces; (2) a heterogeneity-aware load-balancing estimator that guides resource-efficient configuration search; and (3) state caching mechanisms eliminating redundant latency calculations. Evaluations across GPT3 and Llama3 models of varying scales on both homogeneous and heterogeneous clusters demonstrate that AutoHAAP achievessearch efficiency gains,throughput improvements in homogeneous environments, andthroughput enhancements in heterogeneous setups. These results validate AutoHAAP's effectiveness in distributed LLM training on diverse hardware.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get be84d04e-1d4f-4cc6-9b97-4e6f0c876bc7Related papers
- HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed TrainingGuicheng Qi, Junwei Su, Liqi Yang, Tao Li et al.EuroSys 2026 · 1 citation
- Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic ParallelismZizhao Mo, Jianxiong Liao, Huanle Xu, Zhi Zhou et al.SC 2025 · 3 citations
- FASOP: Fast yet Accurate Automated Search for Optimal Parallelization of Transformers on Heterogeneous GPU ClustersSunyeol Hwang, Eungyeong Lee, Hongseok Oh, Youngmin YiHPDC 2024 · 4 citations
- Metis: Fast Automatic Distributed Training on Heterogeneous GPUsTaegeon Um, Byungsoo Oh, Minyoung Kang, Woo-Yeon Lee et al.USENIX ATC 2024 · 81 citations
- AMP: Automatically Finding Model Parallel Strategies with Heterogeneity AwarenessDacheng Li, Hongyi Wang, Eric P. Xing, Hao ZhangNeurIPS 2022 · 57 citations
