StreamingQA: A Benchmark for Adaptation to New Knowledge over Time in Question Answering Models
Adam Liska, Tomás Kociský, Elena Gribovskaya, Tayfun Terzi, Eren Sezener, Devang Agrawal, Cyprien de Masson d'Autume, Tim Scholtes, Manzil Zaheer, Susannah Young, Ellen Gilsenan-McMahon, Sophia Austin
Abstract
Knowledge and language understanding of models evaluated through question answering (QA) has been usually studied on static snapshots of knowledge, like Wikipedia. However, our world is dynamic, evolves over time, and our models' knowledge becomes outdated. To study how semi-parametric QA models and their underlying parametric language models (LMs) adapt to evolving knowledge, we construct a new large-scale dataset, StreamingQA, with human written and generated questions asked on a given date, to be answered from 14 years of time-stamped news articles. We evaluate our models quarterly as they read new articles not seen in pre-training. We show that parametric models can be updated without full retraining, while avoiding catastrophic forgetting. For semi-parametric models, adding new articles into the search space allows for rapid adaptation, however, models with an outdated underlying LM under-perform those with a retrained LM. For questions about higher-frequency named entities, parametric updates are particularly beneficial. In our dynamic world, the StreamingQA dataset enables a more realistic evaluation of QA models, and our experiments highlight several promising directions for future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers44
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsJian Xie, Kai Zhang, Jiangjie Chen, Renze Lou et al.ICLR 2024 · 294 citations
- SealQA: Raising the Bar for Reasoning in Search-Augmented Language ModelsThinh Pham, Nguyen Phan Nguyen, Pratibha Zunjare, Weiyuan Chen et al.ICLR 2026 · 69 citations
- Mass-Editing Memory in a TransformerKevin Meng, Arnab Sen Sharma, Alex J. Andonian, Yonatan Belinkov et al.ICLR 2023 · 52 citations
- Online Adaptation of Language Models with a Memory of Amortized ContextsJihoon Tack, Jaehyung Kim, Eric Mitchell, Jinwoo Shin et al.NeurIPS 2024 · 46 citations
- TiC-CLIP: Continual Training of CLIP ModelsSaurabh Garg, Mehrdad Farajtabar, Hadi Pouransari, Raviteja Vemulapalli et al.ICLR 2024 · 44 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
- Mind the Gap: Assessing Temporal Generalization in Neural Language ModelsAngeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal et al.NeurIPS 2021 · 315 citations
Related papers
- Towards Continual Knowledge Learning of Language ModelsJoel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin et al.ICLR 2022 · 204 citations
- Exploring the Practicality of Generative Retrieval on Dynamic CorporaChaeeun Kim, Soyoung Yoon, Hyunji Lee, Joel Jang et al.EMNLP 2024 · 1 citation
- TiC-LM: A Web-Scale Benchmark for Time-Continual LLM PretrainingJeffrey Li, Mohammadreza Armandpour, Iman Mirzadeh, Sachin Mehta et al.ACL 2025
- TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language ModelsJoel Jang, Seonghyeon Ye, Changho Lee, Sohee Yang et al.EMNLP 2022 · 42 citations
- Are LLMs Prescient? A Continuous Evaluation using Daily News as the OracleHui Dai, Ryan Teehan, Mengye RenICML 2025
