Lipschitz normalization for self-attention layers with application to graph neural networks
George Dasoulas, Kevin Scaman, Aladin Virmaux
摘要
Attention based neural networks are state of the art in a large range of applications. However, their performance tends to degrade when the number of layers increases. In this work, we show that enforcing Lipschitz continuity by normalizing the attention scores can significantly improve the performance of deep attention models. First, we show that, for deep graph attention networks (GAT), gradient explosion appears during training, leading to poor performance of gradient-based training algorithms. To address this issue, we derive a theoretical analysis of the Lipschitz continuity of attention modules and introduce LipschitzNorm, a simple and parameter-free normalization for self-attention mechanisms that enforces the model to be Lipschitz continuous. We then apply LipschitzNorm to GAT and Graph Transformers and show that their performance is substantially improved in the deep setting (10 to 30 layers). More specifically, we show that a deep GAT model with LipschitzNorm achieves state of the art results for node label prediction tasks that exhibit long-range dependencies, while showing consistent improvements over their unnormalized counterparts in benchmark node classification tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- VQ-GNN: A Universal Framework to Scale up Graph Neural Networks using Vector QuantizationMucong Ding, Kezhi Kong, Jingling Li, Chen Zhu 等NeurIPS 2021 · 被引用 68 次
- On the Robustness of Graph Neural Diffusion to Topology PerturbationsYang Song, Qiyu Kang, Sijie Wang, Kai Zhao 等NeurIPS 2022 · 被引用 48 次
- Universality and Limitations of Prompt TuningYihan Wang, Jatin Chauhan, Wei Wang, Cho-Jui HsiehNeurIPS 2023 · 被引用 48 次
- Pay attention to your loss : understanding misconceptions about Lipschitz neural networksLouis Béthune, Thibaut Boissin, Mathieu Serrurier, Franck Mamalet 等NeurIPS 2022 · 被引用 35 次
- DP-Forward: Fine-tuning and Inference on Language Models with Differential Privacy in Forward PassMinxin Du, Xiang Yue, Sherman S. M. Chow, Tianhao Wang 等CCS 2023 · 被引用 35 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding 等ICML 2020 · 被引用 1,910 次
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 被引用 1,586 次
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 被引用 629 次
相关 Paper
- How Smooth Is Attention?Valérie Castin, Pierre Ablin, Gabriel PeyréICML 2024 · 被引用 35 次
- Improving Breadth-Wise Backpropagation in Graph Neural Networks Helps Learning Long-Range DependenciesDenis Lukovnikov, Asja FischerICML 2021 · 被引用 16 次
- Are More Layers Beneficial to Graph Transformers?Haiteng Zhao, Shuming Ma, Dongdong Zhang, Zhi-Hong Deng 等ICLR 2023 · 被引用 5 次
- Enhancing Node-Level Adversarial Defenses by Lipschitz Regularization of Graph Neural NetworksYaning Jia, Dongmian Zou, Hongfei Wang, Hai JinKDD 2023 · 被引用 11 次
- The Lipschitz Constant of Self-AttentionHyunjik Kim, George Papamakarios, Andriy MnihICML 2021 · 被引用 208 次
