Test-Time Training Provably Improves Transformers as In-context Learners
Halil Alperen Gozeten, Muhammed Emrullah Ildiz, Xuechen Zhang, Mahdi Soltanolkotabi, Marco Mondelli, Samet Oymak
Abstract
Test-time training (TTT) methods explicitly update the weights of a model to adapt to the specific test instance, and they have found success in a variety of settings, including most recently language modeling and reasoning. To demystify this success, we investigate a gradient-based TTT algorithm for in-context learning, where we train a transformer model on the in-context demonstrations provided in the test prompt. Specifically, we provide a comprehensive theoretical characterization of linear transformers when the update rule is a single gradient step. Our theory (i) delineates the role of alignment between pretraining distribution and target task, (ii) demystifies how TTT can alleviate distribution shift, and (iii) quantifies the sample complexity of TTT including how it can significantly reduce the eventual sample size required for in-context learning. As our empirical contribution, we study the benefits of TTT for TabPFN, a tabular foundation model. In line with our theory, we demonstrate that TTT significantly reduces the required sample size for tabular classification (3 to 5 times fewer) unlocking substantial inference efficiency with a negligible training cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d632378e-40d4-4516-afdc-6ed21e7edd83Cited by top-tier papers5
- Learning to Correct: Reinforcement Learning for Multi-Attempt Chain-of-ThoughtMuhammed Emrullah Ildiz, Halil Alperen Gozeten, Ege Onur Taga, Samet OymakICML 2026 · 2 citations
- Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention ModelsChungpa Lee, Jy-yong Sohn, Kangwook LeeICML 2026 · 1 citation
- A Decision-Theoretic View of Test-Time Training: When, How Far, and Which Directions to AdaptTomoya WakayamaICML 2026
- Attention with Trained Embeddings Provably Selects Important TokensDiyuan Wu, Aleksandr Shevchenko, Samet Oymak, Marco MondelliNeurIPS 2025
- DecAEvolve: Decompose, Adapt, and Evolve for Effective LLM-based Scientific Equation DiscoveryPouya Behzadifar, Parshin Shojaee, Sanchit Kabra, Kazem Meidani et al.ICML 2026
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller et al.ICML 2020 · 1,220 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
Related papers
- Mitigating Label Shift in Tabular In-Context Learning via Test-Time Posterior AdjustmentSeunghan LeeICML 2026
- The Surprising Effectiveness of Test-Time Training for Few-Shot LearningEkin Akyürek, Mehul Damani, Adam Zweiger, Linlu Qiu et al.ICML 2025
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- SwiftPFN: Revisiting Row-Wise Attention–Only Tabular Foundation Models with Adaptive Early ExitSi-Yang Liu, Han-Jia YeICML 2026
- Specialization after Generalization: Towards Understanding Test-Time Training in Foundation ModelsJonas Hübotter, Patrik Wolf, Aleksandr Shevchenko, Dennis Jüni et al.ICLR 2026 · 6 citations
