Minimal Variance Sampling with Provable Guarantees for Fast Training of Graph Neural Networks
Weilin Cong, Rana Forsati, Mahmut T. Kandemir, Mehrdad Mahdavi
摘要
Sampling methods (e.g., node-wise, layer-wise, or subgraph) has become an indispensable strategy to speed up training large-scale Graph Neural Networks (GNNs). However, existing sampling methods are mostly based on the graph structural information and ignore the dynamicity of optimization, which leads to high variance in estimating the stochastic gradients. The high variance issue can be very pronounced in extremely large graphs, where it results in slow convergence and poor generalization. In this paper, we theoretically analyze the variance of sampling methods and show that, due to the composite structure of empirical risk, the variance of any sampling method can be decomposed into embedding approximation variance in the forward stage and stochastic gradient variance in the backward stage that necessities mitigating both types of variance to obtain faster convergence rate. We propose a decoupled variance reduction strategy that employs (approximate) gradient information to adaptively sample nodes with minimal variance, and explicitly reduces the variance introduced by embedding approximation. We show theoretically and empirically that the proposed method, even with smaller mini-batch sizes, enjoys a faster convergence rate and entails a better generalization compared to the existing methods. Code is public available at here. 1 CCS CONCEPTS • Computing methodologies → Machine learning; Learning latent representations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Decoupling the Depth and Scope of Graph Neural NetworksHanqing Zeng, Muhan Zhang, Yinglong Xia, Ajitesh Srivastava 等NeurIPS 2021 · 被引用 189 次
- GNNAutoScale: Scalable and Expressive Graph Neural Networks via Historical EmbeddingsMatthias Fey, Jan Eric Lenssen, Frank Weichert, Jure LeskovecICML 2021 · 被引用 149 次
- On Provable Benefits of Depth in Training Graph Convolutional NetworksWeilin Cong, Morteza Ramezani, Mehrdad MahdaviNeurIPS 2021 · 被引用 93 次
- PipeGCN: Efficient Full-Graph Training of Graph Convolutional Networks with Pipelined Feature CommunicationCheng Wan, Youjie Li, Cameron R. Wolfe, Anastasios Kyrillidis 等ICLR 2022 · 被引用 89 次
- Accelerating graph sampling for graph machine learning using GPUsAbhinav Jangda, Sandeep Polisetty, Arjun Guha, Marco SerafiniEuroSys 2021 · 被引用 79 次
它引用的顶会 Paper1
相关 Paper
- Bandit Samplers for Training Graph Neural NetworksZiqi Liu, Zhengwei Wu, Zhiqiang Zhang, Jun Zhou 等NeurIPS 2020 · 被引用 55 次
- GCN meets GPU: Decoupling "When to Sample" from "How to Sample"Morteza Ramezani, Weilin Cong, Mehrdad Mahdavi, Anand Sivasubramaniam 等NeurIPS 2020 · 被引用 37 次
- Layer-Neighbor Sampling - Defusing Neighborhood Explosion in GNNsMuhammed Fatih Balin, Ümit V. ÇatalyürekNeurIPS 2023 · 被引用 37 次
- Global Neighbor Sampling for Mixed CPU-GPU Training on Giant GraphsJialin Dong, Da Zheng, Lin F. Yang, George KarypisKDD 2021 · 被引用 22 次
- Resource-Efficient Training for Large Graph Convolutional Networks with Label-Centric Cumulative SamplingMingkai Lin, Wenzhong Li, Ding Li, Yizhou Chen 等WWW 2022 · 被引用 10 次
