On the Initialization of Graph Neural Networks
Jiahang Li, Yakun Song, Xiang Song, David Wipf
摘要
Graph Neural Networks (GNNs) have displayed considerable promise in graph representation learning across various applications. The core learning process requires the initialization of model weight matrices within each GNN layer, which is typically accomplished via classic initialization methods such as Xavier initialization. However, these methods were originally motivated to stabilize the variance of hidden embeddings and gradients across layers of Feedforward Neural Networks (FNNs) and Convolutional Neural Networks (CNNs) to avoid vanishing gradients and maintain steady information flow. In contrast, within the GNN context classical initializations disregard the impact of the input graph structure and message passing on variance. In this paper, we analyze the variance of forward and backward propagation across GNN layers and show that the variance instability of GNN initializations comes from the combined effect of the activation function, hidden dimension, graph structure and message passing. To better account for these influence factors, we propose a new initialization method for Variance Instability Reduction within GNN Optimization (Virgo), which naturally tends to equate forward and backward variances across successive layers. We conduct comprehensive experiments on 15 datasets to show that Virgo can lead to superior model performance and more stable variance at initialization on node classification, link prediction and graph classification tasks. Codes are in https://github.com/LspongebobJH/virgo_icml2023.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Dynamic Rescaling for Training GNNsNimrah Mustafa, Rebekka BurkholzNeurIPS 2024 · 被引用 4 次
- Beyond ReLU: Bifurcation, Oversmoothing, and Topological PriorsErkan Turan, Gaspard Abel, Maysam Behmanesh, Emery Pierson 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper4
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 被引用 1,717 次
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 被引用 1,586 次
- Old can be Gold: Better Gradient Flow can Make Vanilla-GCNs Great AgainAjay Jaiswal, Peihao Wang, Tianlong Chen, Justin F. Rousseau 等NeurIPS 2022 · 被引用 17 次
相关 Paper
- Orthogonal Graph Neural NetworksKai Guo, Kaixiong Zhou, Xia Hu, Yu Li 等AAAI 2022 · 被引用 41 次
- MLPInit: Embarrassingly Simple GNN Training Acceleration with MLP InitializationXiaotian Han, Tong Zhao, Yozen Liu, Xia Hu 等ICLR 2023 · 被引用 12 次
- Rethinking Semi-Supervised Imbalanced Node Classification from Bias-Variance DecompositionDivin Yan, Gengchen Wei, Chen Yang, Shengzhong Zhang 等NeurIPS 2023 · 被引用 27 次
- Inflation Improves Graph Neural NetworksDongxiao He, Rui Guo, Xiaobao Wang, Di Jin 等WWW 2022 · 被引用 5 次
- Elastic Graph Neural NetworksXiaorui Liu, Wei Jin, Yao Ma, Yaxin Li 等ICML 2021 · 被引用 128 次
