Mind the Gap: Assessing Temporal Generalization in Neural Language Models
Angeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d'Autume, Tomás Kociský, Sebastian Ruder, Dani Yogatama, Kris Cao
摘要
Our world is open-ended, non-stationary, and constantly evolving; thus what we talk about and how we talk about it change over time. This inherent dynamic nature of language contrasts with the current static language modelling paradigm, which trains and evaluates models on utterances from overlapping time periods. Despite impressive recent progress, we demonstrate that Transformer-XL language models perform worse in the realistic setup of predicting future utterances from beyond their training period, and that model performance becomes increasingly worse with time. We find that, while increasing model size alone-a key driver behind recent progress-does not solve this problem, having models that continually update their knowledge with new information can indeed mitigate this performance degradation over time. Hence, given the compilation of ever-larger language modelling datasets, combined with the growing list of language-model-based NLP applications that require up-to-date factual knowledge about the world, we argue that now is the right time to rethink the static way in which we currently train and evaluate our language models, and develop adaptive language models that can remain up-to-date with respect to our ever-changing and non-stationary world. We will publicly release our dynamic, streaming language modelling benchmarks for WMT and ARXIV to facilitate language model evaluation that takes temporal dynamics into account. 1 * Equal contribution. ♠ Project initiation. Paper writing. ♦ Project technical infrastructure. ♥ Model design and experiments. ♣ Project support and advice. 1 We release our dynamic (streaming) language modelling benchmark for WMT and ARXIV at https: //github.com/deepmind/deepmind-research/tree/master/pitfalls_static_language_models . 2 In the case of GPT-3 (Brown et al., 2020), such tasks include LAMBADA (Paperno et al., 2016 ), TriviaQA (Joshi et al., 2017b), and WMT translation datasets, among others. These tasks were introduced between 2014 and 2017, which overlap in time with the GPT-3 CommonCrawl dataset that covered the period of 2016-2019.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper75
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn 等ICLR 2022 · 被引用 527 次
- Memory-Based Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning 等ICML 2022 · 被引用 520 次
- Aging with GRACE: Lifelong Model Editing with Discrete Key-Value AdaptorsTom Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim 等NeurIPS 2023 · 被引用 349 次
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsJian Xie, Kai Zhang, Jiangjie Chen, Renze Lou 等ICLR 2024 · 被引用 294 次
- Patching open-vocabulary models by interpolating weightsGabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song 等NeurIPS 2022 · 被引用 230 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
相关 Paper
- StreamingQA: A Benchmark for Adaptation to New Knowledge over Time in Question Answering ModelsAdam Liska, Tomás Kociský, Elena Gribovskaya, Tayfun Terzi 等ICML 2022 · 被引用 129 次
- TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language ModelsJoel Jang, Seonghyeon Ye, Changho Lee, Sohee Yang 等EMNLP 2022 · 被引用 42 次
- Are LLMs Prescient? A Continuous Evaluation using Daily News as the OracleHui Dai, Ryan Teehan, Mengye RenICML 2025
- Improving Temporal Generalization of Pre-trained Language Models with Lexical Semantic ChangeZhaochen Su, Zecheng Tang, Xinyan Guan, Lijun Wu 等EMNLP 2022 · 被引用 11 次
- NEO-BENCH: Evaluating Robustness of Large Language Models with NeologismsJonathan Zheng, Alan Ritter, Wei XuACL 2024
