How Attentive are Graph Attention Networks?
Shaked Brody, Uri Alon, Eran Yahav
摘要
Graph Attention Networks (GATs) are one of the most popular GNN architectures and are considered as the state-of-the-art architecture for representation learning with graphs. In GAT, every node attends to its neighbors given its own representation as the query. However, in this paper we show that GAT computes a very limited kind of attention: the ranking of the attention scores is unconditioned on the query node. We formally define this restricted kind of attention as static attention and distinguish it from a strictly more expressive dynamic attention. Because GATs use a static attention mechanism, there are simple graph problems that GAT cannot express: in a controlled problem, we show that static attention hinders GAT from even fitting the training data. To remove this limitation, we introduce a simple fix by modifying the order of operations and propose GATv2: a dynamic graph attention variant that is strictly more expressive than GAT. We perform an extensive evaluation and show that GATv2 outperforms GAT across 12 OGB and other benchmarks while we match their parametric costs. Our code is available at https://github.com/tech-srl/how_attentive_are_ gats . 1 GATv2 is available as part of the PyTorch Geometric library, 2 the Deep Graph Library, 3 and the TensorFlow GNN library. 4 1 An annotated implementation of GATv2 is available at https://nn.labml.ai/graphs/gatv2/ 2 from torch_geometric.nn.conv.gatv2_conv import GATv2Conv 3 from dgl.nn.pytorch import GATv2Conv 4 from tensorflow_gnn.graph.keras.layers.gat_v2 import GATv2Convolution
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper246
- Pure Transformers are Powerful Graph LearnersJinwoo Kim, Dat Nguyen, Seonwoo Min, Sungjun Cho 等NeurIPS 2022 · 被引用 311 次
- Learning Dynamic Graph Representation of Brain Connectome with Spatio-Temporal AttentionByung-Hoon Kim, Jong Chul Ye, Jae-Jin KimNeurIPS 2021 · 被引用 224 次
- Graph Inductive Biases in Transformers without Message PassingLiheng Ma, Chen Lin, Derek Lim, Adriana Romero-Soriano 等ICML 2023 · 被引用 185 次
- Periodic Graph Transformers for Crystal Material Property PredictionKeqiang Yan, Yi Liu, Yuchao Lin, Shuiwang JiNeurIPS 2022 · 被引用 167 次
- Causal Attention for Interpretable and Generalizable Graph ClassificationYongduo Sui, Xiang Wang, Jiancan Wu, Min Lin 等KDD 2022 · 被引用 166 次
它引用的顶会 Paper13
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- DropEdge: Towards Deep Graph Convolutional Networks on Node ClassificationYu Rong, Wenbing Huang, Tingyang Xu, Junzhou HuangICLR 2020 · 被引用 1,599 次
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 被引用 1,586 次
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan 等ICLR 2020 · 被引用 1,155 次
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie 等NeurIPS 2020 · 被引用 1,113 次
相关 Paper
- HONGAT: Graph Attention Networks in the Presence of High-Order NeighborsHeng-Kai Zhang, Yi-Ge Zhang, Zhi Zhou, Yufeng LiAAAI 2024 · 被引用 15 次
- Learnable Graph Convolutional Attention NetworksAdrián Javaloy, Pablo Sánchez-Martín, Amit Levi, Isabel ValeraICLR 2023 · 被引用 6 次
- GATE: How to Keep Out Intrusive NeighborsNimrah Mustafa, Rebekka BurkholzICML 2024 · 被引用 3 次
- A Dynamical Systems-Inspired Pruning Strategy for Addressing Oversmoothing in Graph Attention NetworksBiswadeep Chakraborty, Harshit Kumar, Saibal MukhopadhyayICML 2025
- Graph External Attention Enhanced TransformerJianqing Liang, Min Chen, Jiye LiangICML 2024 · 被引用 11 次
