Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
Dake Bu, Wei Huang, Andi Han, Atsushi Nitanda, Taiji Suzuki, Qingfu Zhang, Hau-San Wong
Abstract
Transformer-based large language models (LLMs) have displayed remarkable creative prowess and emergence capabilities. Existing empirical studies have revealed a strong connection between these LLMs'impressive emergence abilities and their in-context learning (ICL) capacity, allowing them to solve new tasks using only task-specific prompts without further fine-tuning. On the other hand, existing empirical and theoretical studies also show that there is a linear regularity of the multi-concept encoded semantic representation behind transformer-based LLMs. However, existing theoretical work fail to build up an understanding of the connection between this regularity and the innovative power of ICL. Additionally, prior work often focuses on simplified, unrealistic scenarios involving linear transformers or unrealistic loss functions, and they achieve only linear or sub-linear convergence rates. In contrast, this work provides a fine-grained mathematical analysis to show how transformers leverage the multi-concept semantics of words to enable powerful ICL and excellent out-of-distribution ICL abilities, offering insights into how transformers innovate solutions for certain unseen tasks encoded with multiple cross-concept semantics. Inspired by empirical studies on the linear latent geometry of LLMs, the analysis is based on a concept-based low-noise sparse coding prompt model. Leveraging advanced techniques, this work showcases the exponential 0-1 loss convergence over the highly non-convex training dynamics, which pioneeringly incorporates the challenges of softmax self-attention, ReLU-activated MLPs, and cross-entropy loss. Empirical simulations corroborate the theoretical findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cca76614-d926-45cd-af55-6b6bcfcd2eb5Cited by top-tier papers7
- Trained Mamba Emulates Online Gradient Descent in In-Context Linear RegressionJiarui Jiang, Wei Huang, Miao Zhang, Taiji Suzuki et al.NeurIPS 2025 · 2 citations
- Optimality and NP-Hardness of Transformers in Learning Markovian Dynamical FunctionsYanna Ding, Songtao Lu, Yingdong Lu, Tomasz Nowicki et al.NeurIPS 2025 · 1 citation
- Provable In-Context Vector Arithmetic via Retrieving Task ConceptsDake Bu, Wei Huang, Andi Han, Atsushi Nitanda et al.ICML 2025
- Towards a Theoretical Understanding of In-context Learning: Stability and Non-I.I.D GeneralisationYingjie Wang, Yutian Zhou, Shi Fu, Yuzhu Chen et al.ICLR 2026
- On the Role of Label Noise in the Feature Learning ProcessAndi Han, Wei Huang, Zhanpeng Zhou, Gang Niu et al.ICML 2025
Builds on47
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong et al.NeurIPS 2023 · 356 citations
Related papers
- In-Context Learning with Representations: Contextual Generalization of Trained TransformersTong Yang, Yu Huang, Yingbin Liang, Yuejie ChiNeurIPS 2024 · 45 citations
- Towards More Unified In-Context Visual UnderstandingDianmo Sheng, Dongdong Chen, Zhentao Tan, Qiankun Liu et al.CVPR 2024 · 9 citations
- How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?Hongkang Li, Meng Wang, Songtao Lu, Xiaodong Cui et al.ICML 2024 · 37 citations
- Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in SuperpositionZheyang Xiong, Ziyang Cai, John Cooper, Albert Ge et al.ICML 2025
- From Unstructured Data to In-Context Learning: Exploring What Tasks Can Be Learned and WhenKevin Christian Wibisono, Yixin WangNeurIPS 2024 · 5 citations
