SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding Prediction
Guofan Yu, Sitian Chen, Zhenheng Tang, Xiaowen Chu, Amelie Chi Zhou
摘要
Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges. We present SNI-GNN, a SmartNIC-assisted full-graph training system that reduces communication while preserving accuracy by predicting remote embeddings in-network. SNI-GNN deploys a lightweight linear-trend predictor on SmartNICs to refine cached historical embeddings, coupled with an importance-based boundary-node sampling policy and an asynchronous DPU--GPU data pipeline with intermediate-result reuse. We provide error and convergence bounds showing that predictor bias remains controlled under bounded second-order dynamics and yields standard non-convex convergence with inexact gradients. Implemented on NVIDIA BlueField-3, SNI-GNN integrates with state-of-the-art full-graph systems, cuts communication by 21--45%, achieves 1.3--3.6 end-to-end speedups over BNS-GCN and up to 1.29 over baseline SANCUS, with accuracy loss , and scales efficiently to 16 GPUs on graphs with up to tens of millions of edges. These results indicate SmartNIC-based in-network prediction is a practical complement to partitioning and compression techniques for communication-efficient full-graph GNN training at scale.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Towards Deeper Graph Neural NetworksMeng Liu, Hongyang Gao, Shuiwang JiKDD 2020 · 被引用 496 次
- P3: Distributed Deep Graph Learning at ScaleSwapnil Gandhi, Anand Padmanabha IyerOSDI 2021 · 被引用 192 次
- GNNAutoScale: Scalable and Expressive Graph Neural Networks via Historical EmbeddingsMatthias Fey, Jan Eric Lenssen, Frank Weichert, Jure LeskovecICML 2021 · 被引用 149 次
- A Multi-Scale Approach for Graph Link PredictionLei Cai, Shuiwang JiAAAI 2020 · 被引用 111 次
相关 Paper
- Mithril: A Scalable System for Deep GNN TrainingJingji Chen, Zhuoming Chen, Xuehai QianHPCA 2025 · 被引用 1 次
- Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AIMikhail Khalilov, Salvatore Di Girolamo, Marcin Chrapek, Rami Nudelman 等SC 2024 · 被引用 15 次
- NeutronTask: Scalable and Efficient Multi-GPU GNN Training with Task ParallelismZhenbo Fu, Xin Ai, Qiange Wang, Yanfeng Zhang 等VLDB 2025 · 被引用 4 次
- SC-GNN: A Communication-Efficient Semantic Compression for Distributed Training of GNNsJihe Wang, Ying Wu, Danghui WangDAC 2024 · 被引用 2 次
- ParGNN: A Scalable Graph Neural Network Training Framework on multi-GPUsJunyu Gu, Shunde Li, Rongqiang Cao, Jue Wang 等DAC 2025
