Fundamental limits of learning in sequence multi-index models and deep attention networks: high-dimensional asymptotics and sharp thresholds
Emanuele Troiani, Hugo Cui, Yatin Dandi, Florent Krzakala, Lenka Zdeborová
Abstract
In this manuscript, we study the learning of deep attention neural networks, defined as the composition of multiple self-attention layers, with tied and low-rank weights. We first establish a mapping of such models to sequence multi-index models, a generalization of the widely studied multi-index model to sequential covariates, for which we establish a number of general results. In the context of Bayesian-optimal learning, in the limit of large dimension D and commensurably large number of samples N , we derive a sharp asymptotic characterization of the optimal performance as well as the performance of the best-known polynomial-time algorithm for this setting -namely approximate message-passing-, and characterize sharp thresholds on the minimal sample complexity required for better-than-random prediction performance. Our analysis uncovers, in particular, how the different layers are learned sequentially. Finally, we discuss how this sequential learning can also be observed in a realistic setup.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f1bcb7f-2035-448f-bc83-6a50624f9e29Cited by top-tier papers6
- The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic NetworksVittorio Erba, Emanuele Troiani, Lenka Zdeborová, Florent KrzakalaNeurIPS 2025 · 13 citations
- Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention NetworksLuca Arnaboldi, Bruno Loureiro, Ludovic Stephan, Florent Krzakala et al.NeurIPS 2025 · 10 citations
- Statistical Advantage of Softmax Attention: Insights from Single-Location RegressionO. Duranthon, Pierre Marion, Claire Boyer, Bruno Loureiro et al.ICLR 2026 · 7 citations
- Bayes optimal learning of attention-indexed modelsFabrizio Boncoraglio, Emanuele Troiani, Vittorio Erba, Lenka ZdeborováNeurIPS 2025 · 5 citations
- Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling LawsFabrizio Boncoraglio, Vittorio Erba, Emanuele Troiani, Yizhou Xu et al.ICML 2026 · 5 citations
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Transformers learn to implement preconditioned gradient descent for in-context learningKwangjun Ahn, Xiang Cheng, Hadi Daneshmand, Suvrit SraNeurIPS 2023 · 324 citations
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade et al.NeurIPS 2022 · 220 citations
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang et al.NeurIPS 2022 · 173 citations
Related papers
- Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of TransformersLorenzo Tiberi, Francesca Mignacco, Kazuki Irie, Haim SompolinskyNeurIPS 2024 · 12 citations
- Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic LimitBohan Zhang, Zihao Wang, Hengyu Fu, Jason D. LeeICLR 2026 · 3 citations
- Matrix Inference and Estimation in Multi-Layer ModelsParthe Pandit, Mojtaba Sahraee-Ardakan, Sundeep Rangan, Philip Schniter et al.NeurIPS 2020 · 9 citations
- Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processesSebastian W. Ober, Laurence AitchisonICML 2021 · 65 citations
- Bayes-optimal Learning of Deep Random Networks of Extensive-widthHugo Cui, Florent Krzakala, Lenka ZdeborováICML 2023 · 49 citations
