Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of Transformers
Lorenzo Tiberi, Francesca Mignacco, Kazuki Irie, Haim Sompolinsky
Abstract
Despite the remarkable empirical performance of Transformers, their theoretical understanding remains elusive. Here, we consider a deep multi-head self-attention network, that is closely related to Transformers yet analytically tractable. We develop a statistical mechanics theory of Bayesian learning in this model, deriving exact equations for the network's predictor statistics under the finite-width thermodynamic limit, i.e., , , where is the network width and is the number of training examples. Our theory shows that the predictor statistics are expressed as a sum of independent kernels, each one pairing different 'attention paths', defined as information pathways through different attention heads across layers. The kernels are weighted according to a 'task-relevant kernel combination' mechanism that aligns the total kernel with the task labels. As a consequence, this interplay between attention paths enhances generalization performance. Experiments confirm our findings on both synthetic and real-world sequence classification tasks. Finally, our theory explicitly relates the kernel combination mechanism to properties of the learned weights, allowing for a qualitative transfer of its insights to models trained via gradient descent. As an illustration, we demonstrate an efficient size reduction of the network, by pruning those attention heads that are deemed less relevant by our theory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- High-Dimensional Analysis of Single-Layer Attention for Sparse-Token ClassificationNicholas Barnfield, Hugo Cui, Yue M. LuICLR 2026 · 8 citations
- Perceptrons and Localization of Attention’s Mean-Field LandscapeAntonio Álvarez López, Borjan Geshkovski, Domènec Ruiz-BaletICML 2026 · 7 citations
- Fundamental limits of learning in sequence multi-index models and deep attention networks: high-dimensional asymptotics and sharp thresholdsEmanuele Troiani, Hugo Cui, Yatin Dandi, Florent Krzakala et al.ICML 2025
- When narrower is better: the narrow width limit of Bayesian parallel branching neural networksZechen Zhang, Haim SompolinskyICLR 2025
- Attention with Trained Embeddings Provably Selects Important TokensDiyuan Wu, Aleksandr Shevchenko, Samet Oymak, Marco MondelliNeurIPS 2025
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi et al.ICLR 2020 · 481 citations
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong et al.NeurIPS 2023 · 356 citations
- Deja Vu: Contextual Sparsity for Efficient LLMs at Inference TimeZichang Liu, Jue Wang, Tri Dao, Tianyi Zhou et al.ICML 2023 · 318 citations
Related papers
- Bayes optimal learning of attention-indexed modelsFabrizio Boncoraglio, Emanuele Troiani, Vittorio Erba, Lenka ZdeborováNeurIPS 2025 · 5 citations
- What can a Single Attention Layer Learn? A Study Through the Random Features LensHengyu Fu, Tianyu Guo, Yu Bai, Song MeiNeurIPS 2023 · 47 citations
- Transformer Learns Optimal Variable Selection in Group-Sparse ClassificationChenyang Zhang, Xuran Meng, Yuan CaoICLR 2025
- The Information Pathways Hypothesis: Transformers are Dynamic Self-EnsemblesMd. Shamim Hussain, Mohammed J. Zaki, Dharmashankar SubramanianKDD 2023 · 1 citation
- A Primal-Dual Framework for Transformers and Neural NetworksTan Minh Nguyen, Tam Minh Nguyen, Nhat Ho, Andrea L. Bertozzi et al.ICLR 2023 · 2 citations
