How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?
Jingfeng Wu, Difan Zou, Zixiang Chen, Vladimir Braverman, Quanquan Gu, Peter L. Bartlett
Abstract
Transformers pretrained on diverse tasks exhibit remarkable in-context learning (ICL) capabilities, enabling them to solve unseen tasks solely based on input contexts without adjusting model parameters. In this paper, we study ICL in one of its simplest setups: pretraining a linearly parameterized single-layer linear attention model for linear regression with a Gaussian prior. We establish a statistical task complexity bound for the attention model pretraining, showing that effective pretraining only requires a small number of independent tasks. Furthermore, we prove that the pretrained model closely matches the Bayes optimal algorithm, i.e., optimally tuned ridge regression, by achieving nearly Bayes optimal risk on unseen tasks under a fixed context length. These theoretical findings complement prior experimental research and shed light on the statistical foundations of ICL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c518f31-855d-49e9-8cb9-4eb4ccd76caeCited by top-tier papers60
- The Evolution of Statistical Induction Heads: In-Context Learning Markov ChainsEzra Edelman, Nikolaos Tsilivis, Benjamin L. Edelman, Eran Malach et al.NeurIPS 2024 · 140 citations
- Can large language models explore in-context?Akshay Krishnamurthy, Keegan Harris, Dylan J. Foster, Cyril Zhang et al.NeurIPS 2024 · 95 citations
- A Theoretical Understanding of Self-Correction through In-context AlignmentYifei Wang, Yuyang Wu, Zeming Wei, Stefanie Jegelka et al.NeurIPS 2024 · 69 citations
- Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In ContextXiang Cheng, Yuxin Chen, Suvrit SraICML 2024 · 64 citations
- Why Larger Language Models Do In-context Learning Differently?Zhenmei Shi, Junyi Wei, Zhuoyan Xu, Yingyu LiangICML 2024 · 54 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong et al.NeurIPS 2023 · 356 citations
Related papers
- Pretraining task diversity and the emergence of non-Bayesian in-context learning for regressionAllan Raventós, Mansheej Paul, Feng Chen, Surya GanguliNeurIPS 2023 · 174 citations
- When can in-context learning generalize out of task distribution?Page C. Goddard, Lindsay M. Smith, Vudtiwat Ngampruetikorn, David J. SchwabICML 2025
- In-Context Learning with Representations: Contextual Generalization of Trained TransformersTong Yang, Yu Huang, Yingbin Liang, Yuejie ChiNeurIPS 2024 · 45 citations
- In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-LearningTomoya Wakayama, Taiji SuzukiICML 2026 · 12 citations
- Transformers are Minimax Optimal Nonparametric In-Context LearnersJuno Kim, Tai Nakamaki, Taiji SuzukiNeurIPS 2024 · 42 citations
