Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
Jonas Hübotter, Patrik Wolf, Aleksandr Shevchenko, Dennis Jüni, Andreas Krause, Gil Kur
摘要
Recent empirical studies have explored the idea of continuing to train a model at test-time for a given task, known as test-time training (TTT), and have found it to yield significant performance improvements. However, there is limited understanding of why and when TTT is effective. Earlier explanations mostly focused on the observation that TTT may help when applied to out-of-distribution adaptation or used with privileged data. However, the growing scale of foundation models with most test data being in-distribution questions these explanations. We instead posit that foundation models remain globally underparameterized, with TTT providing a mechanism for specialization after generalization—focusing capacity on concepts relevant to the test task. Specifically, under the linear representation hypothesis, we propose a model in which TTT achieves a substantially smaller in-distribution test error than global training. We empirically validate our model's key assumptions by training a sparse autoencoder on ImageNet, showing that semantically related data points are explained by only a few shared concepts. Finally, we perform scaling studies across image and language tasks that confirm the practical implications of our model, identifying the regimes where specialization is most effective. Beyond TTT, our results provide additional strong evidence in support of the linear representation hypothesis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- One protein is all you needAnton Bushuiev, Roman Bushuiev, Olga Pimenova, Nikola Zadorozhny 等ICLR 2026 · 被引用 1 次
- A Decision-Theoretic View of Test-Time Training: When, How Far, and Which Directions to AdaptTomoya WakayamaICML 2026
- Reinforcement Learning via Self-DistillationJonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann 等ICML 2026
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
相关 Paper
- Test-Time Training Provably Improves Transformers as In-context LearnersHalil Alperen Gozeten, Muhammed Emrullah Ildiz, Xuechen Zhang, Mahdi Soltanolkotabi 等ICML 2025
- Test-time Offline Reinforcement Learning on Goal-related ExperienceMarco Bagatella, Mert Albaba, Jonas Hübotter, Georg Martius 等ICML 2026 · 被引用 7 次
- In-Place Test-Time TrainingGuhao Feng, Shengjie Luo, Kai Hua, Ge Zhang 等ICLR 2026 · 被引用 15 次
- Overcoming Generic Knowledge Loss with Selective Parameter UpdateWenxuan Zhang, Paul Janson, Rahaf Aljundi, Mohamed ElhoseinyCVPR 2024 · 被引用 5 次
- The Surprising Effectiveness of Test-Time Training for Few-Shot LearningEkin Akyürek, Mehul Damani, Adam Zweiger, Linlu Qiu 等ICML 2025
