Old can be Gold: Better Gradient Flow can Make Vanilla-GCNs Great Again
Ajay Jaiswal, Peihao Wang, Tianlong Chen, Justin F. Rousseau, Ying Ding, Zhangyang Wang
摘要
Despite the enormous success of Graph Convolutional Networks (GCNs) in modeling graph-structured data, most of the current GCNs are shallow due to the notoriously challenging problems of over-smoothening and information squashing along with conventional difficulty caused by vanishing gradients and over-fitting. Previous works have been primarily focused on the study of over-smoothening and over-squashing phenomena in training deep GCNs. Surprisingly, in comparison with CNNs/RNNs, very limited attention has been given to understanding how healthy gradient flow can benefit the trainability of deep GCNs. In this paper, firstly, we provide a new perspective of gradient flow to understand the substandard performance of deep GCNs and hypothesize that by facilitating healthy gradient flow, we can significantly improve their trainability, as well as achieve state-of-the-art (SOTA) level performance from vanilla-GCNs [1] . Next, we argue that blindly adopting the Glorot initialization for GCNs is not optimal, and derive a topologyaware isometric initialization scheme for vanilla-GCNs based on the principles of isometry. Additionally, contrary to ad-hoc addition of skip-connections, we propose to use gradient-guided dynamic rewiring of vanilla-GCNs with skip connections. Our dynamic rewiring method uses the gradient flow within each layer during training to introduce on-demand skip-connections adaptively. We provide extensive empirical evidence across multiple datasets that our methods improve gradient flow in deep vanilla-GCNs and significantly boost their performance to comfortably compete and outperform many fancy state-of-the-art methods. Codes are available at: https://github.com/VITA-Group/GradientGCN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Are GATs Out of Balance?Nimrah Mustafa, Aleksandar Bojchevski, Rebekka BurkholzNeurIPS 2023 · 被引用 11 次
- On the Initialization of Graph Neural NetworksJiahang Li, Yakun Song, Xiang Song, David WipfICML 2023 · 被引用 10 次
- Graph Ladling: Shockingly Simple Parallel GNN Training without Intermediate CommunicationAjay Kumar Jaiswal, Shiwei Liu, Tianlong Chen, Ying Ding 等ICML 2023 · 被引用 8 次
- Dynamic Rescaling for Training GNNsNimrah Mustafa, Rebekka BurkholzNeurIPS 2024 · 被引用 4 次
- Customizing Graph Neural Network for CAD Assembly RecommendationFengqi Liang, Huan Zhao, Yuhan Quan, Wei Fang 等KDD 2024 · 被引用 3 次
它引用的顶会 Paper18
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding 等ICML 2020 · 被引用 1,910 次
- DropEdge: Towards Deep Graph Convolutional Networks on Node ClassificationYu Rong, Wenbing Huang, Tingyang Xu, Junzhou HuangICLR 2020 · 被引用 1,599 次
- Geom-GCN: Geometric Graph Convolutional NetworksHongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei 等ICLR 2020 · 被引用 1,445 次
- PairNorm: Tackling Oversmoothing in GNNsLingxiao Zhao, Leman AkogluICLR 2020 · 被引用 590 次
- Towards Deeper Graph Neural NetworksMeng Liu, Hongyang Gao, Shuiwang JiKDD 2020 · 被引用 496 次
相关 Paper
- On Vanishing Gradients, Over-Smoothing, and Over-Squashing in GNNs: Bridging Recurrent and Graph LearningAlvaro Arroyo, Alessio Gravina, Benjamin Gutteridge, Federico Barbero 等NeurIPS 2025 · 被引用 58 次
- DRGCN: Dynamic Evolving Initial Residual for Deep Graph Convolutional NetworksLei Zhang, Xiaodong Yan, Jianshan He, Ruopeng Li 等AAAI 2023 · 被引用 17 次
- Optimization of Graph Neural Networks: Implicit Acceleration by Skip Connections and More DepthKeyulu Xu, Mozhi Zhang, Stefanie Jegelka, Kenji KawaguchiICML 2021 · 被引用 87 次
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 被引用 1,586 次
- DeGNN: Improving Graph Neural Networks with Graph DecompositionXupeng Miao, Nezihe Merve Gürel, Wentao Zhang, Zhichao Han 等KDD 2021 · 被引用 22 次
