In-Context Deep Learning via Transformer Models
Weimin Wu, Maojiang Su, Jerry Yao-Chieh Hu, Zhao Song, Han Liu
Abstract
We investigate the transformer's capability to simulate the training process of deep models via incontext learning (ICL), i.e., in-context deep learning. Our key contribution is providing a positive example of using a transformer to train a deep neural network by gradient descent in an implicit fashion via ICL. Specifically, we provide an explicit construction of a (2N +4)L-layer transformer capable of simulating L gradient descent steps of an N -layer ReLU network through ICL. We also give the theoretical guarantees for the approximation within any given error and the convergence of the ICL gradient descent. Additionally, we extend our analysis to the more practical setting using Softmax-based transformers. We validate our findings on synthetic datasets for 3-layer, 4layer, and 6-layer neural networks. The results show that ICL performance matches that of direct training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 922ae616-e6d1-4030-bf7e-395bd49807cbCited by top-tier papers4
- In-Context Algorithm Emulation in Fixed-Weight TransformersJerry Yao-Chieh Hu, Hude Liu, Jennifer Yuntong Zhang, Han LiuICLR 2026 · 7 citations
- In-Context Universal Approximation, Compositional Generalization, and Algorithm EmulationJerry Yao-Chieh Hu, Hong-Yu Chen, Po-Chiao Lin, Maojiang Su et al.ICML 2026
- In-Context Learning as Conditioned Associative Memory RetrievalWeimin Wu, Teng-Yun Hsiao, Jerry Yao-Chieh Hu, Wenxin Zhang et al.ICML 2025
- How Does the Pretraining Distribution Shape In-Context Learning? A Fundamental Trade-OffWaïss Azizian, Ali HasanICML 2026
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
Related papers
- Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient DescentChenyang Zhang, Yuan CaoICML 2026 · 1 citation
- Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear RegressionDeqing Fu, Tianqi Chen, Robin Jia, Vatsal SharanNeurIPS 2024 · 54 citations
- What learning algorithm is in-context learning? Investigations with linear modelsEkin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma et al.ICLR 2023 · 85 citations
- Chain-of-Thought Gradient DescentHong-Yu Chen, Venkat Ganti, Hude Liu, Jerry Yao-Chieh Hu et al.ICML 2026 · 36 citations
- Transformers are almost optimal metalearners for linear classificationRoey Magen, Gal VardiNeurIPS 2025 · 2 citations
