NeutronStar: Distributed GNN Training with Hybrid Dependency Management
Qiange Wang, Yanfeng Zhang, Hao Wang, Chaoyi Chen, Xiaodong Zhang, Ge Yu
摘要
GNN's training needs to resolve issues of vertex dependencies, i.e., each vertex representation's update depends on its neighbors. Existing distributed GNN systems adopt either a dependencies-cached approach or a dependencies-communicated approach. Having made intensive experiments and analysis, we find that a decision to choose one or the other approach for the best performance is determined by a set of factors, including graph inputs, model configurations, and an underlying computing cluster environment. If various GNN trainings are supported solely by one approach, the performance results are often suboptimal. We study related factors for each GNN training before its execution to choose the best-fit approach accordingly. We propose a hybrid dependency-handling approach that adaptively takes the merits of the two approaches at runtime. Based on the hybrid approach, we further develop a distributed GNN training system called NeutronStar, which makes high performance GNN trainings in an automatic way. NeutronStar is also empowered by effective optimizations in CPU-GPU computation and data processing. Our experimental results on 16-node Aliyun cluster demonstrate that NeutronStar achieves 1.81X-14.25X speedup over existing GNN systems including DistDGL and ROC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Scalable and Efficient Full-Graph GNN Training for Large GraphsXinchen Wan, Kaiqiang Xu, Xudong Liao, Yilun Jin 等SIGMOD 2023 · 被引用 52 次
- Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage AccessesJeongmin Brian Park, Vikram Sharma Mailthody, Zaid Qureshi, Wen-Mei HwuVLDB 2024 · 被引用 37 次
- HongTu: Scalable Full-Graph GNN Training on Multiple GPUsQiange Wang, Yao Chen, Weng-Fai Wong, Bingsheng HeSIGMOD 2024 · 被引用 24 次
- Graph Neural Network Training Systems: A Performance Comparison of Full-Graph and Mini-BatchSaurabh Bajaj, Hui Guan, Marco Serafini, Juelin Liu 等VLDB 2025 · 被引用 19 次
- DAHA: Accelerating GNN Training with Data and Hardware Aware Execution PlanningZhiyuan Li, Xun Jian, Yue Wang, Yingxia Shao 等VLDB 2024 · 被引用 18 次
它引用的顶会 Paper10
- Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless ThreadsJohn Thorpe, Yifan Qiao, Jonathan Eyolfson, Shen Teng 等OSDI 2021 · 被引用 175 次
- GNNAdvisor: An Adaptive and Efficient Runtime System for GNN Acceleration on GPUsYuke Wang, Boyuan Feng, Gushu Li, Shuangchen Li 等OSDI 2021 · 被引用 163 次
- DistGNN: scalable distributed training for large-scale graph neural networksMd. Vasimuddin, Sanchit Misra, Guixiang Ma, Ramanarayan Mohanty 等SC 2021 · 被引用 110 次
- DGCL: an efficient communication library for distributed GNN trainingZhenkun Cai, Xiao Yan, Yidi Wu, Kaihao Ma 等EuroSys 2021 · 被引用 103 次
- Large Graph Convolutional Network Training with GPU-Oriented Data Communication ArchitectureSeungwon Min, Kun Wu, Sitao Huang, Mert Hidayetoglu 等VLDB 2021 · 被引用 85 次
相关 Paper
- NeutronCloud: Resource-Aware Distributed GNN Training in Fluctuating Cloud EnvironmentsMingyi Cao, Chunyu Cao, Yanfeng Zhang, Zhenbo Fu 等VLDB 2026
- NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor ParallelismXin Ai, Hao Yuan, Zeyu Ling, Qiange Wang 等VLDB 2025 · 被引用 8 次
- NeutronOrch: Rethinking Sample-based GNN Training under CPU-GPU Heterogeneous EnvironmentsXin Ai, Qiange Wang, Chunyu Cao, Yanfeng Zhang 等VLDB 2024 · 被引用 19 次
- NeutronHeter: Optimizing Distributed Graph Neural Network Training for Heterogeneous ClustersChunyu Cao, Xin Ai, Qiange Wang, Yanfeng Zhang 等SIGMOD 2026 · 被引用 3 次
- NeutronTask: Scalable and Efficient Multi-GPU GNN Training with Task ParallelismZhenbo Fu, Xin Ai, Qiange Wang, Yanfeng Zhang 等VLDB 2025 · 被引用 4 次
