Test-Time Training on Nearest Neighbors for Large Language Models
Moritz Hardt, Yu Sun
Abstract
Many recent efforts augment language models with retrieval, by adding retrieved data to the input context. For this approach to succeed, the retrieved data must be added at both training and test time. Moreover, as input length grows linearly with the size of retrieved data, cost in computation and memory grows quadratically for modern Transformers. To avoid these complications, we simply fine-tune the model on retrieved data at test time, using its standard training setup. We build a large-scale distributed index based on text embeddings of the Pile dataset. For each test input, our system retrieves its neighbors and fine-tunes the model on their text. Surprisingly, retrieving and training on as few as 20 neighbors, each for only one gradient iteration, drastically improves performance across more than 20 language modeling tasks in the Pile. For example, test-time training with nearest neighbors significantly narrows the performance gap between a small GPT-2 and a GPT-Neo model more than 10 times larger. Sufficient index quality and size, however, are necessary. Our work establishes a first baseline of test-time training for language modeling. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8641b4ea-0eee-44ca-8f66-79dd5f9c34ffCited by top-tier papers48
- TTRL: Test-Time Reinforcement LearningYuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu et al.NeurIPS 2025 · 249 citations
- Learning to Discover at Test TimeMert Yuksekgonul, Daniel Koceja, Xinhao Li, Federico Bianchi et al.ICML 2026 · 73 citations
- Frustratingly Easy Test-Time Adaptation of Vision-Language ModelsMatteo Farina, Gianni Franchi, Giovanni Iacca, Massimiliano Mancini et al.NeurIPS 2024 · 47 citations
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 26 citations
- Transductive Active Learning: Theory and ApplicationsJonas Hübotter, Bhavya Sukhija, Lenart Treven, Yarden As et al.NeurIPS 2024 · 24 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller et al.ICML 2020 · 1,220 citations
Related papers
- Efficiently Learning at Test-Time: Active Fine-Tuning of LLMsJonas Hübotter, Sascha Bongni, Ido Hakimi, Andreas KrauseICLR 2025
- The Surprising Effectiveness of Test-Time Training for Few-Shot LearningEkin Akyürek, Mehul Damani, Adam Zweiger, Linlu Qiu et al.ICML 2025
- Memorizing TransformersYuhuai Wu, Markus Norman Rabe, DeLesley Hutchins, Christian SzegedyICLR 2022 · 231 citations
- Why do Nearest Neighbor Language Models Work?Frank F. Xu, Uri Alon, Graham NeubigICML 2023 · 33 citations
- Studying the Role of Input-Neighbor Overlap in Retrieval-Augmented Language Models Training EfficiencyEhsan Doostmohammadi, Marco KuhlmannEMNLP 2025
