NeutronHeter: Optimizing Distributed Graph Neural Network Training for Heterogeneous Clusters
Chunyu Cao, Xin Ai, Qiange Wang, Yanfeng Zhang, Zhenbo Fu, Hao Yuan, Mingyi Cao, Chaoyi Chen, Yingyou Wen, Yu Gu, Ge Yu
摘要
Distributed training is critical for scaling graph neural networks (GNNs) to large graphs. However, existing systems struggle with efficient GNN training in heterogeneous clusters due to two main challenges. First, mapping GNN workloads to heterogeneous clusters involves matching computation and communication heterogeneity while minimizing communication overhead, resulting in a non-trivial multi-constrained multi-way workload mapping problem. Second, the positive correlation between communication and computational workloads of GNN makes workload mapping difficult, especially in heterogeneous clusters with asymmetric communication bandwidth and computational power (e.g., high-end GPUs with low-bandwidth networks). In this paper, we present NeutronHeter, an efficient GNN training system for heterogeneous clusters. First, we adopt a multi-level workload mapping framework. It converts the original multi-constrained multi-way workload mapping into a more efficient top-down workload mapping on a tree-like resource graph, which is constructed using hierarchical clustering based on computing power and bandwidth. By iteratively mapping the workload layer by layer, our design achieves a near-optimal solution with low overhead. Second, we adopt an adaptive communication migration strategy that reduces communication over slow connections by replicating critical remote vertices on these connections and managing their communication through faster connections. This approach significantly reduces high communication overhead in asymmetric heterogeneous clusters with low-bandwidth links and high computational power. Experimental results in heterogeneous clusters show that NeutronHeter achieves a speedup ranging from 1.06x to 33.05x compared to SOTA frameworks.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- FeLoG: Scalable and Efficient Distributed Graph Embedding with Feedback Loop MechanismPeng Fang, Arijit Khan, Ziqiang Wu, Zhenli Li 等VLDB 2026
- NeutronCloud: Resource-Aware Distributed GNN Training in Fluctuating Cloud EnvironmentsMingyi Cao, Chunyu Cao, Yanfeng Zhang, Zhenbo Fu 等VLDB 2026
相关 Paper
- NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor ParallelismXin Ai, Hao Yuan, Zeyu Ling, Qiange Wang 等VLDB 2025 · 被引用 8 次
- DGCL: an efficient communication library for distributed GNN trainingZhenkun Cai, Xiao Yan, Yidi Wu, Kaihao Ma 等EuroSys 2021 · 被引用 103 次
- NeutronOrch: Rethinking Sample-based GNN Training under CPU-GPU Heterogeneous EnvironmentsXin Ai, Qiange Wang, Chunyu Cao, Yanfeng Zhang 等VLDB 2024 · 被引用 19 次
- NeutronTask: Scalable and Efficient Multi-GPU GNN Training with Task ParallelismZhenbo Fu, Xin Ai, Qiange Wang, Yanfeng Zhang 等VLDB 2025 · 被引用 4 次
- Reducing communication in graph neural network trainingAlok Tripathy, Katherine A. Yelick, Aydin BuluçSC 2020 · 被引用 67 次
