The Lipschitz Constant of Self-Attention
Hyunjik Kim, George Papamakarios, Andriy Mnih
摘要
Lipschitz constants of neural networks have been explored in various contexts in deep learning, such as provable adversarial robustness, estimating Wasserstein distance, stabilising training of GANs, and formulating invertible neural networks. Such works have focused on bounding the Lipschitz constant of fully connected or convolutional networks, composed of linear maps and pointwise non-linearities. In this paper, we investigate the Lipschitz constant of self-attention, a non-linear neural network module widely used in sequence modelling. We prove that the standard dot-product self-attention is not Lipschitz for unbounded input domain, and propose an alternative L2 self-attention that is Lipschitz. We derive an upper bound on the Lipschitz constant of L2 self-attention and provide empirical evidence for its asymptotic tightness. To demonstrate the practical relevance of our theoretical work, we formulate invertible self-attention and use it in a Transformer-based architecture for a characterlevel language modelling task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper84
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun 等NeurIPS 2023 · 被引用 1,178 次
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 被引用 522 次
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 被引用 403 次
- ViTGAN: Training GANs with Vision TransformersKwonjoon Lee, Huiwen Chang, Lu Jiang, Han Zhang 等ICLR 2022 · 被引用 225 次
- Anti-Oversmoothing in Deep Vision Transformers via the Fourier Domain Analysis: From Theory to PracticePeihao Wang, Wenqing Zheng, Tianlong Chen, Zhangyang WangICLR 2022 · 被引用 212 次
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu 等ICML 2020 · 被引用 464 次
- Lipschitz constant estimation of Neural Networks via sparse polynomial optimizationFabian Latorre, Paul Rolland, Volkan CevherICLR 2020 · 被引用 154 次
相关 Paper
- How Smooth Is Attention?Valérie Castin, Pierre Ablin, Gabriel PeyréICML 2024 · 被引用 35 次
- Lipschitz normalization for self-attention layers with application to graph neural networksGeorge Dasoulas, Kevin Scaman, Aladin VirmauxICML 2021 · 被引用 55 次
- Fine-grained Local Sensitivity Analysis of Standard Dot-Product Self-AttentionAaron J. Havens, Alexandre Araujo, Huan Zhang, Bin HuICML 2024 · 被引用 2 次
- Approximation Theory for Lipschitz Continuous TransformersTakashi Furuya, Davide Murari, Carola-Bibiane SchönliebICML 2026 · 被引用 4 次
- LipsFormer: Introducing Lipschitz Continuity to Vision TransformersXianbiao Qi, Jianan Wang, Yihao Chen, Yukai Shi 等ICLR 2023 · 被引用 4 次
