APTMoE: Affinity-Aware Pipeline Tuning for MoE Models on Bandwidth-Constrained GPU Nodes
Yuanxin Wei, Jiangsu Du, Jiazhi Jiang, Xiao Shi, Xianwei Zhang, Dan Huang, Nong Xiao, Yutong Lu
摘要
Recently, the sparsely-gated Mixture-Of-Experts (MoE) architecture has garnered significant attention. To benefit a wider audience, fine-tuning MoE models on more affordable clusters, which are typically a limited number of bandwidthconstrained GPU nodes, holds promise. However, it is non-trivial to apply existing cost-effective fine-tuning approaches to MoE models, due to the increased ratio of data to computation. In this paper, we introduce APTMoE, which employs affinityaware pipeline parallelism for fine-tuning MoE models on bandwidth-constrained GPU nodes. We propose an affinity-aware offloading technique that enhances pipeline parallelism for both computational efficiency and model size, and it benefits from a hierarchical loading strategy and a demand-priority scheduling strategy. To improve the computation efficiency and reduce the data movement volume, the hierarchical loading strategy designs three loading phases and efficiently allocates computation across GPUs and CPUs during these phases, leveraging different levels of expert popularity and computation affinity. With the aim of alleviating the mutual interference among the three loading phases and maximizing the bandwidth utilization, the demand-priority scheduling strategy proactively and dynamically coordinates the loading execution order. Experiments demonstrate that APTMoE outperforms existing methods in most cases. Particularly, APTMoE successfully fine-tunes a 61.2B MoE model on 4 Nvidia A800 GPUs(40GB) and achieves up to throughput improvement compared to the SOTA method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE InferenceShuzhang Zhong, Yanfan Sun, Ling Liang, Runsheng Wang 等DAC 2025 · 被引用 8 次
- Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE ServingYue Pan, Zihan Xia, Po-Kai Hsu, Lanxiang Hu 等MICRO 2025 · 被引用 7 次
- SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State DecouplingAthinagoras Skiadopoulos, Mark Zhao, Swapnil Gandhi, Thomas Norrie 等NSDI 2026 · 被引用 6 次
- Occult: Optimizing Collaborative Communications across Experts for Accelerated Parallel MoE Training and InferenceShuqing Luo, Pingzhi Li, Jie Peng, Yang Zhao 等ICML 2025
- CommitMoE: Efficient Fallback-Free MoE Inference with Offloading Under GPU Memory ConstraintsHan Li, Jingwei Sun, Junqing Lin, Guangzhong SunAAAI 2026
它引用的顶会 Paper12
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase 等USENIX ATC 2021 · 被引用 657 次
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang 等ICML 2022 · 被引用 523 次
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith 等SC 2021 · 被引用 254 次
- Accelerating Distributed MoE Training and Inference with LinaJiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang 等USENIX ATC 2023 · 被引用 191 次
相关 Paper
- Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model TrainingYechan Kim, Hwijoon Lim, Dongsu HanICML 2024 · 被引用 10 次
- Shortcut-connected Expert Parallelism for Accelerating Mixture of ExpertsWeilin Cai, Juyong Jiang, Le Qin, Junwei Cui 等ICML 2025
- Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated SchedulesXinglin Pan, Wenxiang Lin, Shaohuai Shi, Xiaowen Chu 等INFOCOM 2024 · 被引用 13 次
- PipeMoE: Accelerating Mixture-of-Experts through Adaptive PipeliningShaohuai Shi, Xinglin Pan, Xiaowen Chu, Bo LiINFOCOM 2023 · 被引用 23 次
- FlashMoE: Fast Distributed MoE in a Single KernelOsayamen Jonathan Aimuyo, Byungsoo Oh, Rachee SinghNeurIPS 2025 · 被引用 22 次
