Towards Understanding In-Context Learning of Transformers Under Non-I.I.D. Scenarios
Qilu Shen, Yingjie Wang, Jinhai Xiang
Abstract
Understanding the generalization behavior of in-context learning (ICL) in Transformers remains a fundamental challenge, as most existing theoretical analyses are based on the assumption that data are independently and identically distributed (i.i.d.), an assumption that often does not hold in practice. Motivated by the theoretical insight that ICL operates similarly to gradient-based optimization, we leverage the concept of gradient stability to establish generalization error bounds for ICL under a general non-i.i.d. setting. Our analysis shows that two factors play a central role in ICL generalization: the number of demonstrations in the prompt and their distributional alignment with the query. In particular, increasing the number of demonstrations and improving their alignment with the query distribution lead to better generalization, even without any parameter tuning. Under mild conditions, we further prove that the generalization error can achieve the optimal convergence rate of O(N -1 2 ), where N is the number of demonstrations. Our empirical evaluations validate the effectiveness of our theoretical findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1fbb4035-0554-4ac5-9607-d51d0b41adffBuilds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
Related papers
- Towards a Theoretical Understanding of In-context Learning: Stability and Non-I.I.D GeneralisationYingjie Wang, Yutian Zhou, Shi Fu, Yuzhu Chen et al.ICLR 2026
- On the Training Convergence of Transformers for In-Context Classification of Gaussian MixturesWei Shen, Ruida Zhou, Jing Yang, Cong ShenICML 2025
- Transformers as Algorithms: Generalization and Stability in In-context LearningYingcong Li, Muhammed Emrullah Ildiz, Dimitris Papailiopoulos, Samet OymakICML 2023 · 242 citations
- Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear RegressionDeqing Fu, Tianqi Chen, Robin Jia, Vatsal SharanNeurIPS 2024 · 54 citations
- Exact Conversion of In-Context Learning to Model Weights in Linearized-Attention TransformersBrian K. Chen, Tianyang Hu, Hui Jin, Hwee Kuan Lee et al.ICML 2024 · 6 citations
