The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
Ekin Akyürek, Mehul Damani, Adam Zweiger, Linlu Qiu, Han Guo, Jyothish Pari, Yoon Kim, Jacob Andreas
摘要
Language models (LMs) have shown impressive performance on tasks within their training distribution, but often struggle with structurally novel tasks even when given a small number of incontext task examples. We investigate the effectiveness of test-time training (TTT)-temporarily updating model parameters during inference using a loss derived from in-context examplesas a mechanism for improving LMs' reasoning and few-shot learning capabilities. On the Abstraction and Reasoning Corpus (ARC), performing TTT with in-context examples yields up to 6× higher accuracy compared to fine-tuned baselines-reaching 53.0% on the public validation set with an 8B-parameter LM and 61.9% when ensembled with program-synthesis methods, matching average human performance. On BIG-Bench Hard (BBH), TTT on in-context examples surpasses standard few-shot prompting in the 10-shot setting by 7.3 percentage points (50.5% to 57.8%). Our findings highlight the limitations of in-context learning for novel tasks and demonstrate the potential of test-time training to enhance language model adaptability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- TTRL: Test-Time Reinforcement LearningYuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu 等NeurIPS 2025 · 被引用 249 次
- Nested Learning: The Illusion of Deep Learning ArchitecturesAli Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 被引用 96 次
- Self-Adapting Language ModelsAdam Zweiger, Jyothish Pari, Han Guo, Yoon Kim 等NeurIPS 2025 · 被引用 78 次
- Learning to Discover at Test TimeMert Yuksekgonul, Daniel Koceja, Xinhao Li, Federico Bianchi 等ICML 2026 · 被引用 73 次
- Training a Scientific Reasoning Model for ChemistrySiddharth Narayanan, James D. Braza, Ryan-Rhys Griffiths, Albert Bou 等NeurIPS 2025 · 被引用 62 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
相关 Paper
- Context Tuning for In-Context OptimizationJack Lu, Ryan Teehan, Zhenbang Yang, Mengye RenICML 2026
- In-Place Test-Time TrainingGuhao Feng, Shengjie Luo, Kai Hua, Ge Zhang 等ICLR 2026 · 被引用 15 次
- Test-Time Learning for Large Language ModelsJinwu Hu, Zitian Zhang, Guohao Chen, Xutao Wen 等ICML 2025
- TART: A plug-and-play Transformer module for task-agnostic reasoningKush Bhatia, Avanika Narayan, Christopher De Sa, Christopher RéNeurIPS 2023 · 被引用 19 次
- ThinkSum: Probabilistic reasoning over sets using large language modelsBatu Ozturkler, Nikolay Malkin, Zhen Wang, Nebojsa JojicACL 2023 · 被引用 10 次
