Are GATs Out of Balance?
Nimrah Mustafa, Aleksandar Bojchevski, Rebekka Burkholz
Abstract
While the expressive power and computational capabilities of graph neural networks (GNNs) have been theoretically studied, their optimization and learning dynamics, in general, remain largely unexplored. Our study undertakes the Graph Attention Network (GAT), a popular GNN architecture in which a node's neighborhood aggregation is weighted by parameterized attention coefficients. We derive a conservation law of GAT gradient flow dynamics, which explains why a high portion of parameters in GATs with standard initialization struggle to change during training. This effect is amplified in deeper GATs, which perform significantly worse than their shallow counterparts. To alleviate this problem, we devise an initialization scheme that balances the GAT network. Our approach i) allows more effective propagation of gradients and in turn enables trainability of deeper networks, and ii) attains a considerable speedup in training and convergence time in comparison to the standard initialization. Our main theorem serves as a stepping stone to studying the learning dynamics of positive homogeneous models with attention mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Unitary Convolutions for Learning on Graphs and GroupsBobak T. Kiani, Lukas Fesser, Melanie WeberNeurIPS 2024 · 13 citations
- Faithful and Accurate Self-Attention Attribution for Message Passing Neural Networks via the Computation Tree ViewpointYong-Min Shin, Siqing Li, Xin Cao, Won-Yong ShinAAAI 2025 · 6 citations
- Dynamic Rescaling for Training GNNsNimrah Mustafa, Rebekka BurkholzNeurIPS 2024 · 4 citations
- Memorization in Graph Neural NetworksAdarsh Jamadandi, Jing Xu, Adam Dziedzic, Franziska BoenischNeurIPS 2025 · 3 citations
- GATE: How to Keep Out Intrusive NeighborsNimrah Mustafa, Rebekka BurkholzICML 2024 · 3 citations
Builds on34
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 1,717 citations
- DropEdge: Towards Deep Graph Convolutional Networks on Node ClassificationYu Rong, Wenbing Huang, Tingyang Xu, Junzhou HuangICLR 2020 · 1,599 citations
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 1,586 citations
- Geom-GCN: Geometric Graph Convolutional NetworksHongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei et al.ICLR 2020 · 1,445 citations
- Rumor Detection on Social Media with Bi-Directional Graph Convolutional NetworksTian Bian, Xi Xiao, Tingyang Xu, Peilin Zhao et al.AAAI 2020 · 773 citations
Related papers
- Old can be Gold: Better Gradient Flow can Make Vanilla-GCNs Great AgainAjay Jaiswal, Peihao Wang, Tianlong Chen, Justin F. Rousseau et al.NeurIPS 2022 · 17 citations
- Demystifying Oversmoothing in Attention-Based Graph Neural NetworksXinyi Wu, Amir Ajorlou, Zihui Wu, Ali JadbabaieNeurIPS 2023 · 86 citations
- Optimization of Graph Neural Networks: Implicit Acceleration by Skip Connections and More DepthKeyulu Xu, Mozhi Zhang, Stefanie Jegelka, Kenji KawaguchiICML 2021 · 87 citations
- Lipschitz normalization for self-attention layers with application to graph neural networksGeorge Dasoulas, Kevin Scaman, Aladin VirmauxICML 2021 · 55 citations
- Transformative or Conservative? Conservation laws for ResNets and TransformersSibylle Marcotte, Rémi Gribonval, Gabriel PeyréICML 2025
