Lipschitz normalization for self-attention layers with application to graph neural networks
George Dasoulas, Kevin Scaman, Aladin Virmaux
Abstract
Attention based neural networks are state of the art in a large range of applications. However, their performance tends to degrade when the number of layers increases. In this work, we show that enforcing Lipschitz continuity by normalizing the attention scores can significantly improve the performance of deep attention models. First, we show that, for deep graph attention networks (GAT), gradient explosion appears during training, leading to poor performance of gradient-based training algorithms. To address this issue, we derive a theoretical analysis of the Lipschitz continuity of attention modules and introduce LipschitzNorm, a simple and parameter-free normalization for self-attention mechanisms that enforces the model to be Lipschitz continuous. We then apply LipschitzNorm to GAT and Graph Transformers and show that their performance is substantially improved in the deep setting (10 to 30 layers). More specifically, we show that a deep GAT model with LipschitzNorm achieves state of the art results for node label prediction tasks that exhibit long-range dependencies, while showing consistent improvements over their unnormalized counterparts in benchmark node classification tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43235ded-aff1-46a4-b8de-8e325ade6f7bCited by top-tier papers24
- VQ-GNN: A Universal Framework to Scale up Graph Neural Networks using Vector QuantizationMucong Ding, Kezhi Kong, Jingling Li, Chen Zhu et al.NeurIPS 2021 · 68 citations
- On the Robustness of Graph Neural Diffusion to Topology PerturbationsYang Song, Qiyu Kang, Sijie Wang, Kai Zhao et al.NeurIPS 2022 · 48 citations
- Universality and Limitations of Prompt TuningYihan Wang, Jatin Chauhan, Wei Wang, Cho-Jui HsiehNeurIPS 2023 · 48 citations
- Pay attention to your loss : understanding misconceptions about Lipschitz neural networksLouis Béthune, Thibaut Boissin, Mathieu Serrurier, Franck Mamalet et al.NeurIPS 2022 · 35 citations
- DP-Forward: Fine-tuning and Inference on Language Models with Differential Privacy in Forward PassMinxin Du, Xiang Yue, Sherman S. M. Chow, Tianhao Wang et al.CCS 2023 · 35 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding et al.ICML 2020 · 1,910 citations
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 1,586 citations
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
Related papers
- How Smooth Is Attention?Valérie Castin, Pierre Ablin, Gabriel PeyréICML 2024 · 35 citations
- Improving Breadth-Wise Backpropagation in Graph Neural Networks Helps Learning Long-Range DependenciesDenis Lukovnikov, Asja FischerICML 2021 · 16 citations
- Are More Layers Beneficial to Graph Transformers?Haiteng Zhao, Shuming Ma, Dongdong Zhang, Zhi-Hong Deng et al.ICLR 2023 · 5 citations
- Enhancing Node-Level Adversarial Defenses by Lipschitz Regularization of Graph Neural NetworksYaning Jia, Dongmian Zou, Hongfei Wang, Hai JinKDD 2023 · 11 citations
- The Lipschitz Constant of Self-AttentionHyunjik Kim, George Papamakarios, Andriy MnihICML 2021 · 208 citations
