On learning linear dynamical systems in context with attention layers
Maria-Luiza Vladarean, Xuhui Zhang, Suvrit Sra
Abstract
This paper studies the expressive power of linear attention layers for in-context learning (ICL) of linear dynamical systems (LDS). We consider training on sequences of inexact observations produced by noise-corrupted LDSs, with all perturbations being Gaussian. Importantly, this non-i.i.d. data setting is a significant step towards modeling real-world scenarios. We provide the optimal weight construction for a single linear-attention layer and show its equivalence to one step of Gradient Descent relative to an autoregression objective of window size one. Guided by experiments, we uncover a connection to a generalization of the Preconditioned Conjugate Gradient method for larger window sizes. We back our findings with numerical evidence. These results add to the existing understanding of transformers' expressivity as in-context learners and offer plausible hypotheses for recent observations that place their performance on par with that of the Kalman Filter --- the optimal model-dependent learner for this setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
Related papers
- One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-AttentionArvind V. Mahankali, Tatsunori Hashimoto, Tengyu MaICLR 2024 · 160 citations
- In-context Learning of Linear Dynamical Systems with Transformers: Approximation Bounds and Depth-separationFrank Cole, Yuxuan Zhao, Yulong Lu, Tianhao ZhangNeurIPS 2025 · 1 citation
- Linear Transformers are Versatile In-Context LearnersMax Vladymyrov, Johannes von Oswald, Mark Sandler, Rong GeNeurIPS 2024 · 37 citations
- In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD InitializationRuiqi Zhang, Jingfeng Wu, Peter L. BartlettNeurIPS 2024 · 37 citations
- How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?Jingfeng Wu, Difan Zou, Zixiang Chen, Vladimir Braverman et al.ICLR 2024 · 94 citations
