NeutronCloud: Resource-Aware Distributed GNN Training in Fluctuating Cloud Environments
Mingyi Cao, Chunyu Cao, Yanfeng Zhang, Zhenbo Fu, Xin Ai, Qiange Wang, Yu Gu, Ge Yu
Abstract
Graph Neural Networks (GNNs) are widely employed to learn representations from graph-structured data. To support large-scale graph training, researchers use distributed techniques, partitioning the graph across multiple computing nodes and performing parallel training by exchanging dependency vertex information via cross-node communication. However, existing GNN training systems operate on statically partitioned subgraphs, making them difficult to adapt to resource fluctuations. In practice, resource fluctuations in cloud environments often cause variability in compute and communication resources, posing challenges for aligning each worker's workload to its available resources during GNN training. In this paper, we propose NeutronCloud, a system designed for efficient GNN training in cloud environments. First, we adopt a resource-aware workload adjustment strategy. It builds on hybrid dependency handling by obtaining dependency information through both local computation and remote communication. During training, it dynamically adjusts the ratio between locally computed and remotely fetched dependencies based on each worker's available resources, ensuring workload-resource alignment. Second, we employ a dependency-aware partial-reduce approach reusing historical vertex embeddings and skipping the stragglers during gradient aggregation to address extreme resource fluctuations that cause some workers to lag significantly behind others in the cluster. Experimental results on the resource-fluctuating environment demonstrate that NeutronCloud achieves 1.83×-4.43× speedup compared to state-of-the-art distributed GNN systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8adf5fc4-1e79-46f9-ad0e-be40f05a808dBuilds on20
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan et al.ICLR 2020 · 1,155 citations
- Bamboo: Making Preemptible Instances Resilient for Affordable Training of Large DNNsJohn Thorpe, Pengzhan Zhao, Jonathan Eyolfson, Yifan Qiao et al.NSDI 2023 · 144 citations
- DistGNN: scalable distributed training for large-scale graph neural networksMd. Vasimuddin, Sanchit Misra, Guixiang Ma, Ramanarayan Mohanty et al.SC 2021 · 110 citations
- ByteGNN: Efficient Graph Neural Network Training at Large ScaleChenguang Zheng, Hongzhi Chen, Yuxuan Cheng, Zhezheng Song et al.VLDB 2022 · 107 citations
Related papers
- NeutronHeter: Optimizing Distributed Graph Neural Network Training for Heterogeneous ClustersChunyu Cao, Xin Ai, Qiange Wang, Yanfeng Zhang et al.SIGMOD 2026 · 3 citations
- NeutronStar: Distributed GNN Training with Hybrid Dependency ManagementQiange Wang, Yanfeng Zhang, Hao Wang, Chaoyi Chen et al.SIGMOD 2022 · 60 citations
- NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor ParallelismXin Ai, Hao Yuan, Zeyu Ling, Qiange Wang et al.VLDB 2025 · 8 citations
- NeutronTask: Scalable and Efficient Multi-GPU GNN Training with Task ParallelismZhenbo Fu, Xin Ai, Qiange Wang, Yanfeng Zhang et al.VLDB 2025 · 4 citations
- NeutronOrch: Rethinking Sample-based GNN Training under CPU-GPU Heterogeneous EnvironmentsXin Ai, Qiange Wang, Chunyu Cao, Yanfeng Zhang et al.VLDB 2024 · 19 citations
