Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
Deqing Fu, Tianqi Chen, Robin Jia, Vatsal Sharan
Abstract
Transformers excel at in-context learning (ICL) -- learning from demonstrations without parameter updates -- but how they do so remains a mystery. Recent work suggests that Transformers may internally run Gradient Descent (GD), a first-order optimization method, to perform ICL. In this paper, we instead demonstrate that Transformers learn to approximate second-order optimization methods for ICL. For in-context linear regression, Transformers share a similar convergence rate as Iterative Newton's Method, both exponentially faster than GD. Empirically, predictions from successive Transformer layers closely match different iterations of Newton's Method linearly, with each middle layer roughly computing 3 iterations; thus, Transformers and Newton's method converge at roughly the same rate. In contrast, Gradient Descent converges exponentially more slowly. We also show that Transformers can learn in-context on ill-conditioned data, a setting where Gradient Descent struggles but Iterative Newton succeeds. Finally, to corroborate our empirical findings, we prove that Transformers can implement iterations of Newton's method with layers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e63fc508-1aa4-43fa-a54a-9af1090053bcCited by top-tier papers19
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value HeadsZhoutong Wu, Yuan Zhang, Yiming Dong, Chenheng Zhang et al.NeurIPS 2025 · 4 citations
- 'Oh LLM, I'm Asking Thee, Please Give Me a Decision Tree': Zero-Shot Decision Tree Induction and Embedding with Large Language ModelsRicardo Knauer, Mario Koddenbrock, Raphael Wallsberger, Nicholas M. Brisson et al.KDD 2025 · 3 citations
- Linear Transformers Implicitly Discover Unified Numerical AlgorithmsPatrick Lutz, Aditya Gangrade, Hadi Daneshmand, Venkatesh SaligramaNeurIPS 2025 · 3 citations
- Transformers Learn Latent Mixture Models In-Context via Mirror DescentFrancesco D'Angelo, Nicolas FlammarionICLR 2026 · 2 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
Related papers
- On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse RecoveryRenpu Liu, Ruida Zhou, Cong Shen, Jing YangICLR 2025
- In-Context Deep Learning via Transformer ModelsWeimin Wu, Maojiang Su, Jerry Yao-Chieh Hu, Zhao Song et al.ICML 2025
- What learning algorithm is in-context learning? Investigations with linear modelsEkin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma et al.ICLR 2023 · 85 citations
- Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient DescentChenyang Zhang, Yuan CaoICML 2026 · 1 citation
- Towards Understanding In-Context Learning of Transformers Under Non-I.I.D. ScenariosQilu Shen, Yingjie Wang, Jinhai XiangAAAI 2026
