Lune

ICML2024顶会

How Smooth Is Attention?

Valérie Castin, Pierre Ablin, Gabriel Peyré

2024年份
35被引次数
20顶会引用

摘要

Self-attention and masked self-attention are at the heart of Transformers' outstanding success. Still, our mathematical understanding of attention, in particular of its Lipschitz properties - which are key when it comes to analyzing robustness and expressive power - is incomplete. We provide a detailed study of the Lipschitz constant of self-attention in several practical scenarios, discussing the impact of the sequence length nn and layer normalization on the local Lipschitz constant of both unmasked and masked self-attention. In particular, we show that for inputs of length nn in any compact set, the Lipschitz constant of self-attention is bounded by n\sqrt{n} up to a constant factor and that this bound is tight for reasonable sequence lengths. When the sequence length nn is too large for the previous bound to be tight, which we refer to as the mean-field regime, we provide an upper bound and a matching lower bound which are independent of nn. Our mean-field framework for masked self-attention is novel and of independent interest. Our experiments on pretrained and randomly initialized BERT and GPT-2 support our theoretical findings.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper20

问问它们各自怎么用它

它引用的顶会 Paper16

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖