Residual Connections and Normalization Can Provably Prevent Oversmoothing in GNNs
Michael Scholkemper, Xinyi Wu, Ali Jadbabaie, Michael T. Schaub
Abstract
Residual connections and normalization layers have become standard design choices for graph neural networks (GNNs), and were proposed as solutions to the mitigate the oversmoothing problem in GNNs. However, how exactly these methods help alleviate the oversmoothing problem from a theoretical perspective is not well understood. In this work, we provide a formal and precise characterization of (linearized) GNNs with residual connections and normalization layers. We establish that (a) for residual connections, the incorporation of the initial features at each layer can prevent the signal from becoming too smooth, and determines the subspace of possible node representations; (b) batch normalization prevents a complete collapse of the output embedding space to a one-dimensional subspace through the individual rescaling of each column of the feature matrix. This results in the convergence of node representations to the top-k eigenspace of the messagepassing operator; (c) moreover, we show that the centering step of a normalization layer -which can be understood as a projection -alters the graph signal in message-passing in such a way that relevant information can become harder to extract. Building on the last theoretical insight, we introduce GraphNormv2, a novel and principled normalization layer. GraphNormv2 features a learnable centering step designed to preserve the integrity of the original graph signal. Experimental results corroborate the effectiveness of our method, demonstrating improved performance across various GNN architectures and tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02661c46-8fc6-40f4-bf8d-10442de661eeCited by top-tier papers8
- On Vanishing Gradients, Over-Smoothing, and Over-Squashing in GNNs: Bridging Recurrent and Graph LearningAlvaro Arroyo, Alessio Gravina, Benjamin Gutteridge, Federico Barbero et al.NeurIPS 2025 · 58 citations
- Deeper with Riemannian Geometry: Overcoming Oversmoothing and Oversquashing for Graph Foundation ModelsLi Sun, Zhenhao Huang, Ming Zhang, Philip S. YuNeurIPS 2025 · 10 citations
- Are we measuring oversmoothing in graph neural networks correctly?Kaicheng Zhang, Piero Deidda, Desmond Higham, Francesco TudiscoICLR 2026 · 7 citations
- A Signed Graph Approach to Understanding and Mitigating OversmoothingJiaqi Wang, Xinyi Wu, James Cheng, Yifei WangNeurIPS 2025 · 4 citations
- Beyond ReLU: Bifurcation, Oversmoothing, and Topological PriorsErkan Turan, Gaspard Abel, Maysam Behmanesh, Emery Pierson et al.ICML 2026 · 1 citation
Builds on15
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding et al.ICML 2020 · 1,910 citations
- Geom-GCN: Geometric Graph Convolutional NetworksHongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei et al.ICLR 2020 · 1,445 citations
- Graph Neural Networks Exponentially Lose Expressive Power for Node ClassificationKenta Oono, Taiji SuzukiICLR 2020 · 864 citations
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
Related papers
- GraphNorm: A Principled Approach to Accelerating Graph Neural Network TrainingTianle Cai, Shengjie Luo, Keyulu Xu, Di He et al.ICML 2021 · 224 citations
- ContraNorm: A Contrastive Learning Perspective on Oversmoothing and BeyondXiaojun Guo, Yifei Wang, Tianqi Du, Yisen WangICLR 2023 · 10 citations
- PairNorm: Tackling Oversmoothing in GNNsLingxiao Zhao, Leman AkogluICLR 2020 · 590 citations
- Towards Deeper Graph Neural Networks with Differentiable Group NormalizationKaixiong Zhou, Xiao Huang, Yuening Li, Daochen Zha et al.NeurIPS 2020 · 248 citations
- Curriculum-Enhanced Residual Soft An-Isotropic Normalization for Over-Smoothness in Deep GNNsJin Li, Qirong Zhang, Shuling Xu, Xinlong Chen et al.AAAI 2024 · 3 citations
