FlexGNN: A High-Performance, Large-Scale Full-Graph GNN System with Best-Effort Training Plan Optimization
Jeongmin Bae, Donghyoung Han, Min-Soo Kim
Abstract
Recently, full-graph Graph Neural Networks (GNNs) have gained prominence by addressing complex problems such as weather forecasting and material discovery. Existing full-graph training methods do not fully manage intermediate data generated during training and rely on rigid inter-GPU communication, limiting both training speed and scale. We propose FlexGNN, which fully manages intermediate data and adaptively performs inter-GPU communication by generating and optimizing best-effort training execution plans. Extensive experiments demonstrate that FlexGNN significantly outperforms existing full-graph GNN methods in both training speed and scale. Specifically, it is up to 5.4X faster than HongTu and up to 95.5X faster than NeutronStar.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 36c9f3b7-e3b2-44ec-a74a-0c26f213c440Related papers
- HongTu: Scalable Full-Graph GNN Training on Multiple GPUsQiange Wang, Yao Chen, Weng-Fai Wong, Bingsheng HeSIGMOD 2024 · 24 citations
- DistGNN: scalable distributed training for large-scale graph neural networksMd. Vasimuddin, Sanchit Misra, Guixiang Ma, Ramanarayan Mohanty et al.SC 2021 · 110 citations
- Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN TrainingAditya K. Ranjan, Siddharth Singh, Cunyang Wei, Abhinav BhateleSC 2025 · 1 citation
- NeutronHeter: Optimizing Distributed Graph Neural Network Training for Heterogeneous ClustersChunyu Cao, Xin Ai, Qiange Wang, Yanfeng Zhang et al.SIGMOD 2026 · 3 citations
- NeutronCloud: Resource-Aware Distributed GNN Training in Fluctuating Cloud EnvironmentsMingyi Cao, Chunyu Cao, Yanfeng Zhang, Zhenbo Fu et al.VLDB 2026
