Lune

NeurIPS2024Top-tier venue

In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization

Ruiqi Zhang, Jingfeng Wu, Peter L. Bartlett

2024Year
37Citations
18Top-tier citations

Abstract

We study the in-context learning (ICL) ability of a Linear Transformer Block (LTB) that combines a linear attention component and a linear multi-layer perceptron (MLP) component. For ICL of linear regression with a Gaussian prior and a non-zero mean, we show that LTB can achieve nearly Bayes optimal ICL risk. In contrast, using only linear attention must incur an irreducible additive approximation error. Furthermore, we establish a correspondence between LTB and one-step gradient descent estimators with learnable initialization (GD-β\mathsf{GD}\text{-}\mathbf{\beta}), in the sense that every GD-β\mathsf{GD}\text{-}\mathbf{\beta} estimator can be implemented by an LTB estimator and every optimal LTB estimator that minimizes the in-class ICL risk is effectively a GD-β\mathsf{GD}\text{-}\mathbf{\beta} estimator. Finally, we show that GD-β\mathsf{GD}\text{-}\mathbf{\beta} estimators can be efficiently optimized with gradient flow, despite a non-convex training objective. Our results reveal that LTB achieves ICL by implementing GD-β\mathsf{GD}\text{-}\mathbf{\beta}, and they highlight the role of MLP layers in reducing approximation error.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d782e4b0-fedb-4b3f-b190-450fd91cecf4

Cited by top-tier papers18

Ask how each one uses it

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines