Evaluating Temporal Consistency in Multi-Turn Language Models
Yash Kumar Atri, Steven L. Johnson, Thomas Hartvigsen
Abstract
Language models are increasingly deployed in interactive settings where users reason about facts over time rather than in isolation. In such scenarios, correct behavior requires models to maintain and update implicit temporal assumptions established earlier in a conversation. We study this challenge through the lens of temporal scope stability: the ability to preserve, override, or transfer time-scoped factual context across dialogue turns. We introduce ChronoScope, a large-scale diagnostic benchmark designed to isolate temporal scope behavior in controlled multi-turn interactions, comprising over one million deterministically generated question chains grounded in Wikidata. ChronoScope evaluates whether models can correctly retain inferred temporal scope when follow-up questions omit explicit time references, spanning implicit carryover, explicit scope switching, cross-entity transfer, and longer temporal trajectories. Through extensive evaluation of state-of-the-art language models, we find that temporal scope stability is frequently violated in controlled multiturn settings, with models often drifting toward present-day assumptions despite correct underlying knowledge. These failures intensify with interaction length and persist even under oracle context conditions, revealing a gap between single-turn factual accuracy and coherent temporal reasoning under sequential interaction. We make our dataset and evaluation suite publicly available at https://github. com/yashkumaratri/ChronoScope .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f9cbefa-3ffb-47e8-980f-574d4c11fe7dBuilds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LLMs Get Lost In Multi-Turn ConversationPhilippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer NevilleICLR 2026 · 491 citations
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackXingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen et al.ICLR 2024 · 308 citations
- Efficient Sequential Decision Making with Large Language ModelsDingyang Chen, Qi Zhang, Yinglun ZhuEMNLP 2024 · 3 citations
- Promoting Topic Coherence and Inter-Document Consorts in Multi-Document Summarization via Simplicial Complex and Sheaf GraphYash Kumar Atri, Arun Iyer, Tanmoy Chakraborty, Vikram GoyalEMNLP 2023 · 2 citations
Related papers
- From Existence to Exhaustiveness: Unveiling the Compounding Failures of LLMs in Multi-answer Event Temporal ReasoningShaojuan WuSIGIR 2026
- EvolveBench: A Comprehensive Benchmark for Assessing Temporal Awareness in LLMs on Evolving KnowledgeZhiyuan Zhu, Yusheng Liao, Zhe Chen, Yuhao Wang et al.ACL 2025 · 10 citations
- TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language ModelsJoel Jang, Seonghyeon Ye, Changho Lee, Sohee Yang et al.EMNLP 2022 · 42 citations
- ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple DomainsYein Park, Chanwoong Yoon, Jungwoo Park, Donghyeon Lee et al.ICLR 2025
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in LLMsSoyeon Kim, Jindong Wang, Xing Xie, Steven Euijong WhangICLR 2026
