Why do Nearest Neighbor Language Models Work?
Frank F. Xu, Uri Alon, Graham Neubig
摘要
Language models (LMs) compute the probability of a text by sequentially computing a representation of an already-seen context and using this representation to predict the next word. Currently, most LMs calculate these representations through a neural network consuming the immediate previous context. However recently, retrieval-augmented LMs have shown to improve over standard neural LMs, by accessing information retrieved from a large datastore, in addition to their standard, parametric, next-word prediction. In this paper, we set out to understand why retrieval-augmented language models, and specifically why k-nearest neighbor language models (kNN-LMs) perform better than standard parametric LMs, even when the k-nearest neighbor component retrieves examples from the same training set that the LM was originally trained on. To this end, we perform a careful analysis of the various dimensions over which kNN-LM diverges from standard LMs, and investigate these dimensions one by one. Empirically, we identify three main reasons why kNN-LM performs better than standard LMs: using a different input representation for predicting the next tokens, approximate kNN search, and the importance of softmax temperature for the kNN distribution. Further, we incorporate these insights into the model architecture or the training procedure of the standard parametric LM, improving its results without the need for an explicit retrieval component. The code is available at https://github.com/frankxu2004/knnlm-why .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- CAGRA: Highly Parallel Graph Construction and Approximate Nearest Neighbor Search for GPUsHiroyuki Ootomo, Akira Naruse, Corey Nolet, Ray Wang 等ICDE 2024 · 被引用 59 次
- RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation ServingWenqi Jiang, Suvinay Subramanian, Cat Graves, Gustavo Alonso 等ISCA 2025 · 被引用 16 次
- Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language ModelsJiaqi Cao, Jiarui Wang, Rubin Wei, Qipeng Guo 等NeurIPS 2025 · 被引用 14 次
- FT2Ra: A Fine-Tuning-Inspired Approach to Retrieval-Augmented Code CompletionQi Guo, Xiaohong Li, Xiaofei Xie, Shangqing Liu 等ISSTA 2024 · 被引用 14 次
- kNN-LM Does Not Improve Open-ended Text GenerationShufan Wang, Yixiao Song, Andrew Drozdov, Aparna Garimella 等EMNLP 2023 · 被引用 3 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
- Nearest Neighbor Machine TranslationUrvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2021 · 被引用 323 次
相关 Paper
- Nearest Neighbor Zero-Shot InferenceWeijia Shi, Julian Michael, Suchin Gururangan, Luke ZettlemoyerEMNLP 2022 · 被引用 26 次
- Efficient Nearest Neighbor Language ModelsJunxian He, Graham Neubig, Taylor Berg-KirkpatrickEMNLP 2021
- Neuro-Symbolic Language Modeling with Automaton-augmented RetrievalUri Alon, Frank F. Xu, Junxian He, Sudipta Sengupta 等ICML 2022 · 被引用 79 次
- The Inductive Bias of In-Context Learning: Rethinking Pretraining Example DesignYoav Levine, Noam Wies, Daniel Jannai, Dan Navon 等ICLR 2022 · 被引用 43 次
- DReX: Accurate and Scalable Dense Retrieval Acceleration via Algorithmic-Hardware CodesignDerrick Quinn, E. Ezgi Yücel, Martin Prammer, Zhenxing Fan 等ISCA 2025 · 被引用 10 次
