Provably Learning a Multi-head Attention Layer
Sitan Chen, Yuanzhi Li
Abstract
The multi-head attention layer is one of the key components of the transformer architecture that sets it apart from traditional feed-forward models. Given a sequence length k, attention matrices Σ1,…,Σm∈ℝd× d, and projection matrices W1,…,Wm∈ℝd× d, the corresponding multi-head attention layer F: ℝk× d→ ℝk× d transforms length-k sequences of d-dimensional tokens X∈ℝk× d via F(X) ≜ ∑i=1m softmax(XΣiX⊤)XWi. In this work, we initiate the study of provably learning a multi-head attention layer from random examples and give the first nontrivial upper and lower bounds for this problem. Provided Wi, Σi satisfy certain non-degeneracy conditions, we give a (dk)O(m3)-time algorithm that learns F to small error given random labeled examples drawn uniformly from ± 1k× d. We also prove computational lower bounds showing that in the worst case, exponential dependence on the number of heads m is unavoidable. We chose to focus on Boolean X to mimic the discrete nature of tokens in large language models, though our techniques naturally extend to standard continuous settings, e.g. Gaussian. Our algorithm, which is centered around using examples to sculpt a convex body containing the unknown parameters, is a significant departure from existing provable algorithms for learning feed-forward networks, which predominantly exploit fine-grained algebraic and rotation invariance properties of the Gaussian distribution. In contrast, our analysis is more flexible as it primarily relies on various upper and lower tail bounds for the input distribution and “slices” thereof.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d4265dd-0f40-41eb-8542-e7b47db749c8Cited by top-tier papers16
- Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in TransformersSiyu Chen, Heejune Sheen, Tianhao Wang, Zhuoran YangNeurIPS 2024 · 48 citations
- How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?Hongkang Li, Meng Wang, Songtao Lu, Xiaodong Cui et al.ICML 2024 · 37 citations
- Implicit Optimization Bias of Next-token Prediction in Linear ModelsChristos ThrampoulidisNeurIPS 2024 · 19 citations
- Non-asymptotic Convergence of Training Transformers for Next-token PredictionRuiquan Huang, Yingbin Liang, Jing YangNeurIPS 2024 · 15 citations
- Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow AnalysisHongru Yang, Bhavya Kailkhura, Zhangyang Wang, Yingbin LiangNeurIPS 2024 · 14 citations
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi et al.ICLR 2020 · 481 citations
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong et al.NeurIPS 2023 · 356 citations
Related papers
- What can a Single Attention Layer Learn? A Study Through the Random Features LensHengyu Fu, Tianyu Guo, Yu Bai, Song MeiNeurIPS 2023 · 47 citations
- On the Computational Hardness of TransformersBarna Saha, Yinzhan Xu, Christopher Ye, Hantao YuSTOC 2026
- Provably Learning Attention with QueriesSatwik Bhattamishra, Kulin Shah, Michael Hahn, Varun KanadeICML 2026
- Memorization Capacity of Multi-Head Attention in TransformersSadegh Mahdavi, Renjie Liao, Christos ThrampoulidisICLR 2024 · 34 citations
- Theoretical limitations of multi-layer TransformerLijie Chen, Binghui Peng, Hongxun WuFOCS 2025 · 2 citations
