In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization
Ruiqi Zhang, Jingfeng Wu, Peter L. Bartlett
Abstract
We study the in-context learning (ICL) ability of a Linear Transformer Block (LTB) that combines a linear attention component and a linear multi-layer perceptron (MLP) component. For ICL of linear regression with a Gaussian prior and a non-zero mean, we show that LTB can achieve nearly Bayes optimal ICL risk. In contrast, using only linear attention must incur an irreducible additive approximation error. Furthermore, we establish a correspondence between LTB and one-step gradient descent estimators with learnable initialization (), in the sense that every estimator can be implemented by an LTB estimator and every optimal LTB estimator that minimizes the in-class ICL risk is effectively a estimator. Finally, we show that estimators can be efficiently optimized with gradient flow, despite a non-convex training objective. Our results reveal that LTB achieves ICL by implementing , and they highlight the role of MLP layers in reducing approximation error.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d782e4b0-fedb-4b3f-b190-450fd91cecf4Cited by top-tier papers18
- Looped Transformers are Better at Learning Learning AlgorithmsLiu Yang, Kangwook Lee, Robert D. Nowak, Dimitris PapailiopoulosICLR 2024 · 82 citations
- Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention LandscapeJuno Kim, Taiji SuzukiICML 2024 · 42 citations
- Transformers are Minimax Optimal Nonparametric In-Context LearnersJuno Kim, Tai Nakamaki, Taiji SuzukiNeurIPS 2024 · 42 citations
- Pretrained Transformer Efficiently Learns Low-Dimensional Target Functions In-ContextKazusato Oko, Yujin Song, Taiji Suzuki, Denny WuNeurIPS 2024 · 34 citations
- How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear RegressionXingwu Chen, Lei Zhao, Difan ZouNeurIPS 2024 · 19 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong et al.NeurIPS 2023 · 356 citations
Related papers
- How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?Jingfeng Wu, Difan Zou, Zixiang Chen, Vladimir Braverman et al.ICLR 2024 · 94 citations
- On learning linear dynamical systems in context with attention layersMaria-Luiza Vladarean, Xuhui Zhang, Suvrit SraICLR 2026
- Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient DescentChenyang Zhang, Yuan CaoICML 2026 · 1 citation
- Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear RegressionDeqing Fu, Tianqi Chen, Robin Jia, Vatsal SharanNeurIPS 2024 · 54 citations
- One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-AttentionArvind V. Mahankali, Tatsunori Hashimoto, Tengyu MaICLR 2024 · 160 citations
