Minimax Rates for Learning Pairwise Interactions in Attention-Style Models
Shai Zucker, Xiong Wang, Fei Lu, Inbar Seroussi
Abstract
We study the convergence rate of learning pairwise interactions in single-layer attention-style models, where tokens interact through a weight matrix and a nonlinear activation function. We prove that the minimax rate is , where is the sample size and is the Hölder smoothness of the activation function. Importantly, this rate is independent of the embedding dimension , the number of tokens , and the rank of the weight matrix, provided that . These results highlight a fundamental statistical efficiency of attention-style models, even when the weight matrix and activation are not separately identifiable, and provide a theoretical understanding of attention mechanisms and guidance on training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b2e6f90-8853-4499-bdd9-5e0543d93e2dBuilds on16
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- SOFT: Softmax-free Transformer with Linear ComplexityJiachen Lu, Jinghan Yao, Junge Zhang, Xiatian Zhu et al.NeurIPS 2021 · 232 citations
- The emergence of clusters in self-attention dynamicsBorjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, Philippe RigolletNeurIPS 2023 · 163 citations
Related papers
- Token Sample Complexity of AttentionLéa Bohbot, Cyril Letrouit, Gabriel Peyré, François-Xavier VialardICML 2026 · 1 citation
- High-Dimensional Analysis of Single-Layer Attention for Sparse-Token ClassificationNicholas Barnfield, Hugo Cui, Yue M. LuICLR 2026 · 8 citations
- Efficient and Minimax Optimal In-context Nonparametric Regression with TransformersMichelle Ching, Ioana Popescu, Nico Smith, Tianyi Ma et al.ICML 2026 · 6 citations
- Bayes optimal learning of attention-indexed modelsFabrizio Boncoraglio, Emanuele Troiani, Vittorio Erba, Lenka ZdeborováNeurIPS 2025 · 5 citations
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
