MariusGNN: Resource-Efficient Out-of-Core Training of Graph Neural Networks
Roger Waleffe, Jason Mohoney, Theodoros Rekatsinas, Shivaram Venkataraman
摘要
We study training of Graph Neural Networks (GNNs) for large-scale graphs. We revisit the premise of using distributed training for billion-scale graphs and show that for graphs that fit in main memory or the SSD of a single machine, out-of-core pipelined training with a single GPU can outperform state-of-the-art (SoTA) multi-GPU solutions. We introduce MariusGNN, the first system that utilizes the entire storage hierarchy---including disk---for GNN training. MariusGNN introduces a series of data organization and algorithmic contributions that 1) minimize the end-to-end time required for training and 2) ensure that models learned with disk-based training exhibit accuracy similar to those fully trained in memory. We evaluate MariusGNN against SoTA systems for learning GNN models and find that single-GPU training in MariusGNN achieves the same level of accuracy up to 8× faster than multi-GPU training in these systems, thus, introducing an order of magnitude monetary cost reduction. MariusGNN is open-sourced at www.marius-project.org.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- BeaconGNN: Large-Scale GNN Acceleration with Out-of-Order Streaming In-Storage ComputingYuyue Wang, Xiurui Pan, Yuda An, Jie Zhang 等HPCA 2024 · 被引用 27 次
- OUTRE: An OUT-of-core De-REdundancy GNN Training Framework for Massive Graphs within A Single MachineZeang Sheng, Wentao Zhang, Yangyu Tao, Bin CuiVLDB 2024 · 被引用 17 次
- Quake: Adaptive Indexing for Vector SearchJason Mohoney, Devesh Sarda, Mengze Tang, Shihabur Rahman Chowdhury 等OSDI 2025 · 被引用 12 次
- Lotan: Bridging the Gap between GNNs and Scalable Graph Analytics EnginesYuhao Zhang, Arun KumarVLDB 2023 · 被引用 11 次
- Buffalo: Enabling Large-Scale GNN Training via Memory-Efficient BucketizationShuangyan Yang, Minjia Zhang, Dong LiHPCA 2025 · 被引用 10 次
它引用的顶会 Paper14
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan 等ICLR 2020 · 被引用 1,155 次
- Discrete Graph Structure Learning for Forecasting Multiple Time SeriesChao Shang, Jie Chen, Jinbo BiICLR 2021 · 被引用 353 次
- Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless ThreadsJohn Thorpe, Yifan Qiao, Jonathan Eyolfson, Shen Teng 等OSDI 2021 · 被引用 175 次
- GNNAdvisor: An Adaptive and Efficient Runtime System for GNN Acceleration on GPUsYuke Wang, Boyuan Feng, Gushu Li, Shuangchen Li 等OSDI 2021 · 被引用 163 次
相关 Paper
- Marius: Learning Massive Graph Embeddings on a Single MachineJason Mohoney, Roger Waleffe, Henry Xu, Theodoros Rekatsinas 等OSDI 2021 · 被引用 75 次
- DiskGNN: Bridging I/O Efficiency and Model Accuracy for Out-of-Core GNN TrainingRenjie Liu, Yichuan Wang, Xiao Yan, Haitian Jiang 等SIGMOD 2025 · 被引用 8 次
- Ginex: SSD-enabled Billion-scale Graph Neural Network Training on a Single Machine via Provably Optimal In-memory CachingYeonhong Park, Sunhong Min, Jae W. LeeVLDB 2022 · 被引用 57 次
- Capsule: An Out-of-Core Training Mechanism for Colossal GNNsYongan Xiang, Zezhong Ding, Rui Guo, Shangyou Wang 等SIGMOD 2025 · 被引用 6 次
- CaliEX: A Disk-Based Large-Scale GNN Training System with Joint Design of Caching and ExecutionCan Su, Haipeng Zhang, Hanyu Zhao, Wenting Shen 等ICDE 2025 · 被引用 1 次
