On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse Recovery
Renpu Liu, Ruida Zhou, Cong Shen, Jing Yang
Abstract
An intriguing property of the Transformer is its ability to perform in-context learning (ICL), where the Transformer can solve different inference tasks without parameter updating based on the contextual information provided by the corresponding input-output demonstration pairs. It has been theoretically proved that ICL is enabled by the capability of Transformers to perform gradient-descent algorithms (Von Oswald et al., 2023a; Bai et al., 2024) . This work takes a step further and shows that Transformers can perform learning-to-optimize (L2O) algorithms. Specifically, for the ICL sparse recovery (formulated as LASSO) tasks, we show that a K-layer Transformer can perform an L2O algorithm with a provable convergence rate linear in K. This provides a new perspective explaining the superior ICL capability of Transformers, even with only a few layers, which cannot be achieved by the standard gradient-descent algorithms. Moreover, unlike the conventional L2O algorithms that require the measurement matrix involved in training to match that in testing, the trained Transformer is able to solve sparse recovery problems generated with different measurement matrices. Besides, Transformers as an L2O algorithm can leverage structural information embedded in the training tasks to accelerate its convergence during ICL, and generalize across different lengths of demonstration pairs, where conventional L2O algorithms typically struggle or fail. Such theoretical findings are supported by our experimental results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed73f146-6869-4a1c-b00b-25d4aae2cc53Cited by top-tier papers3
- Unlabeled Data Can Provably Enhance In-Context Learning of TransformersRenpu Liu, Jing YangNeurIPS 2025 · 3 citations
- Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear RegressionXingwu Chen, Lu Miao, Beining Wu, Difan ZouICML 2026
- Understanding Generalization and Forgetting in In-Context Continual LearningGuangyu Li, Meng Ding, Lijie HuICML 2026
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
Related papers
- Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear RegressionDeqing Fu, Tianqi Chen, Robin Jia, Vatsal SharanNeurIPS 2024 · 54 citations
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong et al.NeurIPS 2023 · 356 citations
- Towards Understanding In-Context Learning of Transformers Under Non-I.I.D. ScenariosQilu Shen, Yingjie Wang, Jinhai XiangAAAI 2026
- On the Training Convergence of Transformers for In-Context Classification of Gaussian MixturesWei Shen, Ruida Zhou, Jing Yang, Cong ShenICML 2025
- Transformers are almost optimal metalearners for linear classificationRoey Magen, Gal VardiNeurIPS 2025 · 2 citations
