NeutronOrch: Rethinking Sample-based GNN Training under CPU-GPU Heterogeneous Environments
Xin Ai, Qiange Wang, Chunyu Cao, Yanfeng Zhang, Chaoyi Chen, Hao Yuan, Yu Gu, Ge Yu
摘要
Graph Neural Networks (GNNs) have demonstrated outstanding performance in various applications. Existing frameworks utilize CPU-GPU heterogeneous environments to train GNN models and integrate mini-batch and sampling techniques to overcome the GPU memory limitation. In CPU-GPU heterogeneous environments, we can divide sample-based GNN training into three steps: sample, gather, and train. Existing GNN systems use different task orchestrating methods to employ each step on the CPU or GPU. After extensive experiments and analysis, we find that existing task orchestrating methods fail to fully utilize the heterogeneous resources, limited by inefficient CPU processing or GPU resource contention. In this paper, we propose NeutronOrch, a system for samplebased GNN training that incorporates a hotness-aware layer-based task orchestrating method and ensures balanced utilization of the CPU and GPU. NeutronOrch decouples the training process by layer and pushes down the training task of the bottom layer to the CPU. This significantly reduces the computational load and memory footprint of GPU training. To avoid inefficient CPU processing, Neu-tronOrch only offloads the training of frequently accessed vertices to the CPU and lets GPU reuse their embeddings with bounded staleness. Furthermore, NeutronOrch provides a fine-grained pipeline design for the layer-based task orchestrating method, fully overlapping different tasks on heterogeneous resources while strictly guaranteeing bounded staleness. The experimental results show that compared with the state-of-the-art GNN systems, NeutronOrch can achieve up to 11.51× performance speedup.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Graph Neural Network Training Systems: A Performance Comparison of Full-Graph and Mini-BatchSaurabh Bajaj, Hui Guan, Marco Serafini, Juelin Liu 等VLDB 2025 · 被引用 19 次
- NeutronTask: Scalable and Efficient Multi-GPU GNN Training with Task ParallelismZhenbo Fu, Xin Ai, Qiange Wang, Yanfeng Zhang 等VLDB 2025 · 被引用 4 次
- A Comprehensive Benchmark on Spectral GNNs: The Impact on Efficiency, Memory, and EffectivenessNingyi Liao, Haoyu Liu, Zulun Zhu, Siqiang Luo 等SIGMOD 2026 · 被引用 4 次
- Faster Convergence in Mini-batch Graph Neural Networks Training with Pseudo Full Neighborhood CompensationQiqi Zhou, Yanyan Shen, Lei ChenVLDB 2025 · 被引用 2 次
- CLM: Removing the GPU Memory Barrier for 3D Gaussian SplattingHexu Zhao, Xiwen Min, Xiaoteng Liu, Moonjun Gong 等ASPLOS 2026 · 被引用 1 次
它引用的顶会 Paper18
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase 等USENIX ATC 2021 · 被引用 657 次
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith 等SC 2021 · 被引用 254 次
- GNNAutoScale: Scalable and Expressive Graph Neural Networks via Historical EmbeddingsMatthias Fey, Jan Eric Lenssen, Frank Weichert, Jure LeskovecICML 2021 · 被引用 149 次
相关 Paper
- NeutronCloud: Resource-Aware Distributed GNN Training in Fluctuating Cloud EnvironmentsMingyi Cao, Chunyu Cao, Yanfeng Zhang, Zhenbo Fu 等VLDB 2026
- NeutronHeter: Optimizing Distributed Graph Neural Network Training for Heterogeneous ClustersChunyu Cao, Xin Ai, Qiange Wang, Yanfeng Zhang 等SIGMOD 2026 · 被引用 3 次
- NeutronStar: Distributed GNN Training with Hybrid Dependency ManagementQiange Wang, Yanfeng Zhang, Hao Wang, Chaoyi Chen 等SIGMOD 2022 · 被引用 60 次
- FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large ScaleZeyu Zhu, Peisong Wang, Qinghao Hu, Gang Li 等ASPLOS 2024 · 被引用 8 次
- FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network TrainingKezhao Huang, Haitian Jiang, Minjie Wang, Guangxuan Xiao 等VLDB 2024 · 被引用 13 次
