NeutronHeter: Optimizing Distributed Graph Neural Network Training for Heterogeneous Clusters
Chunyu Cao, Xin Ai, Qiange Wang, Yanfeng Zhang, Zhenbo Fu, Hao Yuan, Mingyi Cao, Chaoyi Chen, Yingyou Wen, Yu Gu, Ge Yu
Abstract
Distributed training is critical for scaling graph neural networks (GNNs) to large graphs. However, existing systems struggle with efficient GNN training in heterogeneous clusters due to two main challenges. First, mapping GNN workloads to heterogeneous clusters involves matching computation and communication heterogeneity while minimizing communication overhead, resulting in a non-trivial multi-constrained multi-way workload mapping problem. Second, the positive correlation between communication and computational workloads of GNN makes workload mapping difficult, especially in heterogeneous clusters with asymmetric communication bandwidth and computational power (e.g., high-end GPUs with low-bandwidth networks). In this paper, we present NeutronHeter, an efficient GNN training system for heterogeneous clusters. First, we adopt a multi-level workload mapping framework. It converts the original multi-constrained multi-way workload mapping into a more efficient top-down workload mapping on a tree-like resource graph, which is constructed using hierarchical clustering based on computing power and bandwidth. By iteratively mapping the workload layer by layer, our design achieves a near-optimal solution with low overhead. Second, we adopt an adaptive communication migration strategy that reduces communication over slow connections by replicating critical remote vertices on these connections and managing their communication through faster connections. This approach significantly reduces high communication overhead in asymmetric heterogeneous clusters with low-bandwidth links and high computational power. Experimental results in heterogeneous clusters show that NeutronHeter achieves a speedup ranging from 1.06x to 33.05x compared to SOTA frameworks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6d1f2a43-bd20-43bf-be56-b2d222b385b4Cited by top-tier papers2
- FeLoG: Scalable and Efficient Distributed Graph Embedding with Feedback Loop MechanismPeng Fang, Arijit Khan, Ziqiang Wu, Zhenli Li et al.VLDB 2026
- NeutronCloud: Resource-Aware Distributed GNN Training in Fluctuating Cloud EnvironmentsMingyi Cao, Chunyu Cao, Yanfeng Zhang, Zhenbo Fu et al.VLDB 2026
Related papers
- NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor ParallelismXin Ai, Hao Yuan, Zeyu Ling, Qiange Wang et al.VLDB 2025 · 8 citations
- DGCL: an efficient communication library for distributed GNN trainingZhenkun Cai, Xiao Yan, Yidi Wu, Kaihao Ma et al.EuroSys 2021 · 103 citations
- NeutronOrch: Rethinking Sample-based GNN Training under CPU-GPU Heterogeneous EnvironmentsXin Ai, Qiange Wang, Chunyu Cao, Yanfeng Zhang et al.VLDB 2024 · 19 citations
- NeutronTask: Scalable and Efficient Multi-GPU GNN Training with Task ParallelismZhenbo Fu, Xin Ai, Qiange Wang, Yanfeng Zhang et al.VLDB 2025 · 4 citations
- Reducing communication in graph neural network trainingAlok Tripathy, Katherine A. Yelick, Aydin BuluçSC 2020 · 67 citations
