Are GATs Out of Balance?
Nimrah Mustafa, Aleksandar Bojchevski, Rebekka Burkholz
摘要
While the expressive power and computational capabilities of graph neural networks (GNNs) have been theoretically studied, their optimization and learning dynamics, in general, remain largely unexplored. Our study undertakes the Graph Attention Network (GAT), a popular GNN architecture in which a node's neighborhood aggregation is weighted by parameterized attention coefficients. We derive a conservation law of GAT gradient flow dynamics, which explains why a high portion of parameters in GATs with standard initialization struggle to change during training. This effect is amplified in deeper GATs, which perform significantly worse than their shallow counterparts. To alleviate this problem, we devise an initialization scheme that balances the GAT network. Our approach i) allows more effective propagation of gradients and in turn enables trainability of deeper networks, and ii) attains a considerable speedup in training and convergence time in comparison to the standard initialization. Our main theorem serves as a stepping stone to studying the learning dynamics of positive homogeneous models with attention mechanisms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Unitary Convolutions for Learning on Graphs and GroupsBobak T. Kiani, Lukas Fesser, Melanie WeberNeurIPS 2024 · 被引用 13 次
- Faithful and Accurate Self-Attention Attribution for Message Passing Neural Networks via the Computation Tree ViewpointYong-Min Shin, Siqing Li, Xin Cao, Won-Yong ShinAAAI 2025 · 被引用 6 次
- Dynamic Rescaling for Training GNNsNimrah Mustafa, Rebekka BurkholzNeurIPS 2024 · 被引用 4 次
- Memorization in Graph Neural NetworksAdarsh Jamadandi, Jing Xu, Adam Dziedzic, Franziska BoenischNeurIPS 2025 · 被引用 3 次
- GATE: How to Keep Out Intrusive NeighborsNimrah Mustafa, Rebekka BurkholzICML 2024 · 被引用 3 次
它引用的顶会 Paper34
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 被引用 1,717 次
- DropEdge: Towards Deep Graph Convolutional Networks on Node ClassificationYu Rong, Wenbing Huang, Tingyang Xu, Junzhou HuangICLR 2020 · 被引用 1,599 次
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 被引用 1,586 次
- Geom-GCN: Geometric Graph Convolutional NetworksHongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei 等ICLR 2020 · 被引用 1,445 次
- Rumor Detection on Social Media with Bi-Directional Graph Convolutional NetworksTian Bian, Xi Xiao, Tingyang Xu, Peilin Zhao 等AAAI 2020 · 被引用 773 次
相关 Paper
- Old can be Gold: Better Gradient Flow can Make Vanilla-GCNs Great AgainAjay Jaiswal, Peihao Wang, Tianlong Chen, Justin F. Rousseau 等NeurIPS 2022 · 被引用 17 次
- Demystifying Oversmoothing in Attention-Based Graph Neural NetworksXinyi Wu, Amir Ajorlou, Zihui Wu, Ali JadbabaieNeurIPS 2023 · 被引用 86 次
- Optimization of Graph Neural Networks: Implicit Acceleration by Skip Connections and More DepthKeyulu Xu, Mozhi Zhang, Stefanie Jegelka, Kenji KawaguchiICML 2021 · 被引用 87 次
- Lipschitz normalization for self-attention layers with application to graph neural networksGeorge Dasoulas, Kevin Scaman, Aladin VirmauxICML 2021 · 被引用 55 次
- Transformative or Conservative? Conservation laws for ResNets and TransformersSibylle Marcotte, Rémi Gribonval, Gabriel PeyréICML 2025
