On the Initialization of Graph Neural Networks
Jiahang Li, Yakun Song, Xiang Song, David Wipf
Abstract
Graph Neural Networks (GNNs) have displayed considerable promise in graph representation learning across various applications. The core learning process requires the initialization of model weight matrices within each GNN layer, which is typically accomplished via classic initialization methods such as Xavier initialization. However, these methods were originally motivated to stabilize the variance of hidden embeddings and gradients across layers of Feedforward Neural Networks (FNNs) and Convolutional Neural Networks (CNNs) to avoid vanishing gradients and maintain steady information flow. In contrast, within the GNN context classical initializations disregard the impact of the input graph structure and message passing on variance. In this paper, we analyze the variance of forward and backward propagation across GNN layers and show that the variance instability of GNN initializations comes from the combined effect of the activation function, hidden dimension, graph structure and message passing. To better account for these influence factors, we propose a new initialization method for Variance Instability Reduction within GNN Optimization (Virgo), which naturally tends to equate forward and backward variances across successive layers. We conduct comprehensive experiments on 15 datasets to show that Virgo can lead to superior model performance and more stable variance at initialization on node classification, link prediction and graph classification tasks. Codes are in https://github.com/LspongebobJH/virgo_icml2023.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a91fa20-7268-45d0-af6e-e7b78d8b5738Cited by top-tier papers2
- Dynamic Rescaling for Training GNNsNimrah Mustafa, Rebekka BurkholzNeurIPS 2024 · 4 citations
- Beyond ReLU: Bifurcation, Oversmoothing, and Topological PriorsErkan Turan, Gaspard Abel, Maysam Behmanesh, Emery Pierson et al.ICML 2026 · 1 citation
Builds on4
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 1,717 citations
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 1,586 citations
- Old can be Gold: Better Gradient Flow can Make Vanilla-GCNs Great AgainAjay Jaiswal, Peihao Wang, Tianlong Chen, Justin F. Rousseau et al.NeurIPS 2022 · 17 citations
Related papers
- Orthogonal Graph Neural NetworksKai Guo, Kaixiong Zhou, Xia Hu, Yu Li et al.AAAI 2022 · 41 citations
- MLPInit: Embarrassingly Simple GNN Training Acceleration with MLP InitializationXiaotian Han, Tong Zhao, Yozen Liu, Xia Hu et al.ICLR 2023 · 12 citations
- Rethinking Semi-Supervised Imbalanced Node Classification from Bias-Variance DecompositionDivin Yan, Gengchen Wei, Chen Yang, Shengzhong Zhang et al.NeurIPS 2023 · 27 citations
- Inflation Improves Graph Neural NetworksDongxiao He, Rui Guo, Xiaobao Wang, Di Jin et al.WWW 2022 · 5 citations
- Elastic Graph Neural NetworksXiaorui Liu, Wei Jin, Yao Ma, Yaxin Li et al.ICML 2021 · 128 citations
