Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
Jonas Hübotter, Patrik Wolf, Aleksandr Shevchenko, Dennis Jüni, Andreas Krause, Gil Kur
Abstract
Recent empirical studies have explored the idea of continuing to train a model at test-time for a given task, known as test-time training (TTT), and have found it to yield significant performance improvements. However, there is limited understanding of why and when TTT is effective. Earlier explanations mostly focused on the observation that TTT may help when applied to out-of-distribution adaptation or used with privileged data. However, the growing scale of foundation models with most test data being in-distribution questions these explanations. We instead posit that foundation models remain globally underparameterized, with TTT providing a mechanism for specialization after generalization—focusing capacity on concepts relevant to the test task. Specifically, under the linear representation hypothesis, we propose a model in which TTT achieves a substantially smaller in-distribution test error than global training. We empirically validate our model's key assumptions by training a sparse autoencoder on ImageNet, showing that semantically related data points are explained by only a few shared concepts. Finally, we perform scaling studies across image and language tasks that confirm the practical implications of our model, identifying the regimes where specialization is most effective. Beyond TTT, our results provide additional strong evidence in support of the linear representation hypothesis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e59587c4-6d46-4336-95c8-c7f4ec01af79Cited by top-tier papers3
- One protein is all you needAnton Bushuiev, Roman Bushuiev, Olga Pimenova, Nikola Zadorozhny et al.ICLR 2026 · 1 citation
- A Decision-Theoretic View of Test-Time Training: When, How Far, and Which Directions to AdaptTomoya WakayamaICML 2026
- Reinforcement Learning via Self-DistillationJonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann et al.ICML 2026
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller et al.ICML 2020 · 1,220 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
Related papers
- Test-Time Training Provably Improves Transformers as In-context LearnersHalil Alperen Gozeten, Muhammed Emrullah Ildiz, Xuechen Zhang, Mahdi Soltanolkotabi et al.ICML 2025
- Test-time Offline Reinforcement Learning on Goal-related ExperienceMarco Bagatella, Mert Albaba, Jonas Hübotter, Georg Martius et al.ICML 2026 · 7 citations
- In-Place Test-Time TrainingGuhao Feng, Shengjie Luo, Kai Hua, Ge Zhang et al.ICLR 2026 · 15 citations
- Overcoming Generic Knowledge Loss with Selective Parameter UpdateWenxuan Zhang, Paul Janson, Rahaf Aljundi, Mohamed ElhoseinyCVPR 2024 · 5 citations
- The Surprising Effectiveness of Test-Time Training for Few-Shot LearningEkin Akyürek, Mehul Damani, Adam Zweiger, Linlu Qiu et al.ICML 2025
