Lune

NeurIPS2021顶会

Mind the Gap: Assessing Temporal Generalization in Neural Language Models

Angeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d'Autume, Tomás Kociský, Sebastian Ruder, Dani Yogatama, Kris Cao

2021年份
315被引次数
75顶会引用

摘要

Our world is open-ended, non-stationary, and constantly evolving; thus what we talk about and how we talk about it change over time. This inherent dynamic nature of language contrasts with the current static language modelling paradigm, which trains and evaluates models on utterances from overlapping time periods. Despite impressive recent progress, we demonstrate that Transformer-XL language models perform worse in the realistic setup of predicting future utterances from beyond their training period, and that model performance becomes increasingly worse with time. We find that, while increasing model size alone-a key driver behind recent progress-does not solve this problem, having models that continually update their knowledge with new information can indeed mitigate this performance degradation over time. Hence, given the compilation of ever-larger language modelling datasets, combined with the growing list of language-model-based NLP applications that require up-to-date factual knowledge about the world, we argue that now is the right time to rethink the static way in which we currently train and evaluate our language models, and develop adaptive language models that can remain up-to-date with respect to our ever-changing and non-stationary world. We will publicly release our dynamic, streaming language modelling benchmarks for WMT and ARXIV to facilitate language model evaluation that takes temporal dynamics into account. 1 * Equal contribution. ♠ Project initiation. Paper writing. ♦ Project technical infrastructure. ♥ Model design and experiments. ♣ Project support and advice. 1 We release our dynamic (streaming) language modelling benchmark for WMT and ARXIV at https: //github.com/deepmind/deepmind-research/tree/master/pitfalls_static_language_models . 2 In the case of GPT-3 (Brown et al., 2020), such tasks include LAMBADA (Paperno et al., 2016 ), TriviaQA (Joshi et al., 2017b), and WMT translation datasets, among others. These tasks were introduced between 2014 and 2017, which overlap in time with the GPT-3 CommonCrawl dataset that covered the period of 2016-2019.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper75

问问它们各自怎么用它

它引用的顶会 Paper14

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖