Lune

AAAI2026Top-tier venue

Towards Understanding In-Context Learning of Transformers Under Non-I.I.D. Scenarios

Qilu Shen, Yingjie Wang, Jinhai Xiang

2026Year

Abstract

Understanding the generalization behavior of in-context learning (ICL) in Transformers remains a fundamental challenge, as most existing theoretical analyses are based on the assumption that data are independently and identically distributed (i.i.d.), an assumption that often does not hold in practice. Motivated by the theoretical insight that ICL operates similarly to gradient-based optimization, we leverage the concept of gradient stability to establish generalization error bounds for ICL under a general non-i.i.d. setting. Our analysis shows that two factors play a central role in ICL generalization: the number of demonstrations in the prompt and their distributional alignment with the query. In particular, increasing the number of demonstrations and improving their alignment with the query distribution lead to better generalization, even without any parameter tuning. Under mild conditions, we further prove that the generalization error can achieve the optimal convergence rate of O(N -1 2 ), where N is the number of demonstrations. Our empirical evaluations validate the effectiveness of our theoretical findings.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 1fbb4035-0554-4ac5-9607-d51d0b41adff

Builds on17

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines