Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent
Chenyang Zhang, Yuan Cao
Abstract
Transformers have demonstrated remarkable in-context learning (ICL) capabilities. The strong ICL performance of transformers is commonly believed to arise from their ability to implicitly execute certain algorithms on the context, thereby enhancing prediction and generation. In this work, we investigate how transformers with softmax attention perform in-context learning on linear classification data. We first construct a class of multi-layer transformers that can perform in-context logistic regression, with each layer exactly performing one step of normalized gradient descent on an in-context loss. Then, we show that our constructed transformer can be obtained through (i) training a single self-attention layer supervised by one-step gradient descent, and (ii) recurrently applying the trained layer to obtain a looped model. Training convergence guarantees of the self-attention layer and out-of-distribution generalization guarantees of the looped model are provided. Our results advance the theoretical understanding of ICL mechanism by showcasing how softmax transformers can effectively act as in-context learners.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on44
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
Related papers
- The Closeness of In-Context Learning and Weight Shifting for Softmax RegressionShuai Li, Zhao Song, Yu Xia, Tong Yu et al.NeurIPS 2024 · 53 citations
- In-Context Deep Learning via Transformer ModelsWeimin Wu, Maojiang Su, Jerry Yao-Chieh Hu, Zhao Song et al.ICML 2025
- In-Context Learning with Transformers: Softmax Attention Adapts to Function LipschitznessLiam Collins, Advait Parulekar, Aryan Mokhtari, Sujay Sanghavi et al.NeurIPS 2024 · 33 citations
- One-Layer Transformer Provably Learns One-Nearest Neighbor In ContextZihao Li, Yuan Cao, Cheng Gao, Yihan He et al.NeurIPS 2024 · 25 citations
- In-context Convergence of TransformersYu Huang, Yuan Cheng, Yingbin LiangICML 2024 · 114 citations
