Old can be Gold: Better Gradient Flow can Make Vanilla-GCNs Great Again
Ajay Jaiswal, Peihao Wang, Tianlong Chen, Justin F. Rousseau, Ying Ding, Zhangyang Wang
Abstract
Despite the enormous success of Graph Convolutional Networks (GCNs) in modeling graph-structured data, most of the current GCNs are shallow due to the notoriously challenging problems of over-smoothening and information squashing along with conventional difficulty caused by vanishing gradients and over-fitting. Previous works have been primarily focused on the study of over-smoothening and over-squashing phenomena in training deep GCNs. Surprisingly, in comparison with CNNs/RNNs, very limited attention has been given to understanding how healthy gradient flow can benefit the trainability of deep GCNs. In this paper, firstly, we provide a new perspective of gradient flow to understand the substandard performance of deep GCNs and hypothesize that by facilitating healthy gradient flow, we can significantly improve their trainability, as well as achieve state-of-the-art (SOTA) level performance from vanilla-GCNs [1] . Next, we argue that blindly adopting the Glorot initialization for GCNs is not optimal, and derive a topologyaware isometric initialization scheme for vanilla-GCNs based on the principles of isometry. Additionally, contrary to ad-hoc addition of skip-connections, we propose to use gradient-guided dynamic rewiring of vanilla-GCNs with skip connections. Our dynamic rewiring method uses the gradient flow within each layer during training to introduce on-demand skip-connections adaptively. We provide extensive empirical evidence across multiple datasets that our methods improve gradient flow in deep vanilla-GCNs and significantly boost their performance to comfortably compete and outperform many fancy state-of-the-art methods. Codes are available at: https://github.com/VITA-Group/GradientGCN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e5d5375-1248-4bcf-98ec-9320426003f1Cited by top-tier papers6
- Are GATs Out of Balance?Nimrah Mustafa, Aleksandar Bojchevski, Rebekka BurkholzNeurIPS 2023 · 11 citations
- On the Initialization of Graph Neural NetworksJiahang Li, Yakun Song, Xiang Song, David WipfICML 2023 · 10 citations
- Graph Ladling: Shockingly Simple Parallel GNN Training without Intermediate CommunicationAjay Kumar Jaiswal, Shiwei Liu, Tianlong Chen, Ying Ding et al.ICML 2023 · 8 citations
- Dynamic Rescaling for Training GNNsNimrah Mustafa, Rebekka BurkholzNeurIPS 2024 · 4 citations
- Customizing Graph Neural Network for CAD Assembly RecommendationFengqi Liang, Huan Zhao, Yuhan Quan, Wei Fang et al.KDD 2024 · 3 citations
Builds on18
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding et al.ICML 2020 · 1,910 citations
- DropEdge: Towards Deep Graph Convolutional Networks on Node ClassificationYu Rong, Wenbing Huang, Tingyang Xu, Junzhou HuangICLR 2020 · 1,599 citations
- Geom-GCN: Geometric Graph Convolutional NetworksHongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei et al.ICLR 2020 · 1,445 citations
- PairNorm: Tackling Oversmoothing in GNNsLingxiao Zhao, Leman AkogluICLR 2020 · 590 citations
- Towards Deeper Graph Neural NetworksMeng Liu, Hongyang Gao, Shuiwang JiKDD 2020 · 496 citations
Related papers
- On Vanishing Gradients, Over-Smoothing, and Over-Squashing in GNNs: Bridging Recurrent and Graph LearningAlvaro Arroyo, Alessio Gravina, Benjamin Gutteridge, Federico Barbero et al.NeurIPS 2025 · 58 citations
- DRGCN: Dynamic Evolving Initial Residual for Deep Graph Convolutional NetworksLei Zhang, Xiaodong Yan, Jianshan He, Ruopeng Li et al.AAAI 2023 · 17 citations
- Optimization of Graph Neural Networks: Implicit Acceleration by Skip Connections and More DepthKeyulu Xu, Mozhi Zhang, Stefanie Jegelka, Kenji KawaguchiICML 2021 · 87 citations
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 1,586 citations
- DeGNN: Improving Graph Neural Networks with Graph DecompositionXupeng Miao, Nezihe Merve Gürel, Wentao Zhang, Zhichao Han et al.KDD 2021 · 22 citations
