Lune

ICML2024Top-tier venue

How Smooth Is Attention?

Valérie Castin, Pierre Ablin, Gabriel Peyré

2024Year
35Citations
20Top-tier citations

Abstract

Self-attention and masked self-attention are at the heart of Transformers' outstanding success. Still, our mathematical understanding of attention, in particular of its Lipschitz properties - which are key when it comes to analyzing robustness and expressive power - is incomplete. We provide a detailed study of the Lipschitz constant of self-attention in several practical scenarios, discussing the impact of the sequence length nn and layer normalization on the local Lipschitz constant of both unmasked and masked self-attention. In particular, we show that for inputs of length nn in any compact set, the Lipschitz constant of self-attention is bounded by n\sqrt{n} up to a constant factor and that this bound is tight for reasonable sequence lengths. When the sequence length nn is too large for the previous bound to be tight, which we refer to as the mean-field regime, we provide an upper bound and a matching lower bound which are independent of nn. Our mean-field framework for masked self-attention is novel and of independent interest. Our experiments on pretrained and randomly initialized BERT and GPT-2 support our theoretical findings.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ee66873f-dd70-4afd-8da6-8c0241babdb1

Cited by top-tier papers20

Ask how each one uses it

Builds on16

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines