Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws
Fabrizio Boncoraglio, Vittorio Erba, Emanuele Troiani, Yizhou Xu, FLORENT KRZAKALA, Lenka Zdeborova
Abstract
Trained attention layers exhibit striking and reproducible spectral structure of the weights, including low-rank collapse, bulk deformation, and isolated spectral outliers, yet the origin of these phenomena and their implications for generalization remain poorly understood. We study empirical risk minimization in a single-head tied-attention layer trained on synthetic high-dimensional sequence tasks generated from the attention-indexed model. Using tools from random matrix theory, spin-glass theory, and approximate message passing, we obtain an exact high-dimensional characterization of training and test error, interpolation and recovery thresholds, and the spectrum of the key and query matrices. Our theory predicts the full singular-value distribution of the trained query–key map—including low-rank structure and isolated spectral outliers—in qualitative agreement with observations in more realistic transformers. Finally, for targets with power-law spectra, we show that learning proceeds through sequential spectral recovery, leading to the emergence of power-law scaling laws.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2c33e5f-fab3-4c0c-8b36-865b09e1607dCited by top-tier papers1
Ask how each one uses itBuilds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- An empirical analysis of compute-optimal large language model trainingJordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya et al.NeurIPS 2022 · 566 citations
- Pretraining task diversity and the emergence of non-Bayesian in-context learning for regressionAllan Raventós, Mansheej Paul, Feng Chen, Surya GanguliNeurIPS 2023 · 174 citations
Related papers
- Bayes optimal learning of attention-indexed modelsFabrizio Boncoraglio, Emanuele Troiani, Vittorio Erba, Lenka ZdeborováNeurIPS 2025 · 5 citations
- What can a Single Attention Layer Learn? A Study Through the Random Features LensHengyu Fu, Tianyu Guo, Yu Bai, Song MeiNeurIPS 2023 · 47 citations
- Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning RegimeLeonardo Defilippis, Yizhou Xu, Julius Girardin, Vittorio Erba et al.ICLR 2026 · 20 citations
- Fundamental limits of learning in sequence multi-index models and deep attention networks: high-dimensional asymptotics and sharp thresholdsEmanuele Troiani, Hugo Cui, Yatin Dandi, Florent Krzakala et al.ICML 2025
- On the Role of Attention Masks and LayerNorm in TransformersXinyi Wu, Amir Ajorlou, Yifei Wang, Stefanie Jegelka et al.NeurIPS 2024 · 54 citations
