AMP: Automatically Finding Model Parallel Strategies with Heterogeneity Awareness
Dacheng Li, Hongyi Wang, Eric P. Xing, Hao Zhang
摘要
Scaling up model sizes can lead to fundamentally new capabilities in many machine learning (ML) tasks. However, training big models requires strong distributed system expertise to carefully design model-parallel execution strategies that suit the model architectures and cluster setups. In this paper, we develop AMP, a framework that automatically derives such strategies. AMP identifies a valid space of model parallelism strategies and efficiently searches the space for high-performed strategies, by leveraging a cost model designed to capture the heterogeneity of the model and cluster specifications. Unlike existing methods, AMP is specifically tailored to support complex models composed of uneven layers and cluster setups with more heterogeneous accelerators and bandwidth. We evaluate AMP on popular models and cluster setups from public clouds and show that AMP returns parallel strategies that match the expert-tuned strategies on typical cluster setups. On heterogeneous clusters or models with heterogeneous architectures, AMP finds strategies with 1.54× and 1.77× higher throughput than state-of-the-art model-parallel systems, respectively. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Metis: Fast Automatic Distributed Training on Heterogeneous GPUsTaegeon Um, Byungsoo Oh, Minyoung Kang, Woo-Yeon Lee 等USENIX ATC 2024 · 被引用 81 次
- FlexSP: Accelerating Large Language Model Training via Flexible Sequence ParallelismYujie Wang, Shiju Wang, Shenhan Zhu, Fangcheng Fu 等ASPLOS 2025 · 被引用 8 次
- Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-OptimizationZhanda Zhu, Christina Giannoula, Muralidhar Andoorveedu, Qidong Su 等EuroSys 2025 · 被引用 8 次
- Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data AnnotationsHaoyang Li, Fangcheng Fu, Hao Ge, Sheng Lin 等OSDI 2026 · 被引用 7 次
- Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model ParallelizationHaoyang Li, Fangcheng Fu, Hao Ge, Sheng Lin 等SIGMOD 2025 · 被引用 6 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale UpYifan Jiang, Shiyu Chang, Zhangyang WangNeurIPS 2021 · 被引用 515 次
相关 Paper
- Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningLianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang 等OSDI 2022 · 被引用 75 次
- UniAP: Unifying Inter- and Intra-Layer Automatic Parallelism by Mixed Integer Quadratic ProgrammingHao Lin, Ke Wu, Jie Li, Jun Li 等CVPR 2025
- AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM TrainingYuanyuan Wang, Nana Tang, Yuyang Wang, Shu Pan 等HPCA 2026
- HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed TrainingGuicheng Qi, Junwei Su, Liqi Yang, Tao Li 等EuroSys 2026 · 被引用 1 次
- nnScaler: Constraint-Guided Parallelization Plan Generation for Deep Learning TrainingZhiqi Lin, Youshan Miao, Quanlu Zhang, Fan Yang 等OSDI 2024 · 被引用 38 次
