Evaluating Temporal Consistency in Multi-Turn Language Models
Yash Kumar Atri, Steven L. Johnson, Thomas Hartvigsen
摘要
Language models are increasingly deployed in interactive settings where users reason about facts over time rather than in isolation. In such scenarios, correct behavior requires models to maintain and update implicit temporal assumptions established earlier in a conversation. We study this challenge through the lens of temporal scope stability: the ability to preserve, override, or transfer time-scoped factual context across dialogue turns. We introduce ChronoScope, a large-scale diagnostic benchmark designed to isolate temporal scope behavior in controlled multi-turn interactions, comprising over one million deterministically generated question chains grounded in Wikidata. ChronoScope evaluates whether models can correctly retain inferred temporal scope when follow-up questions omit explicit time references, spanning implicit carryover, explicit scope switching, cross-entity transfer, and longer temporal trajectories. Through extensive evaluation of state-of-the-art language models, we find that temporal scope stability is frequently violated in controlled multiturn settings, with models often drifting toward present-day assumptions despite correct underlying knowledge. These failures intensify with interaction length and persist even under oracle context conditions, revealing a gap between single-turn factual accuracy and coherent temporal reasoning under sequential interaction. We make our dataset and evaluation suite publicly available at https://github. com/yashkumaratri/ChronoScope .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LLMs Get Lost In Multi-Turn ConversationPhilippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer NevilleICLR 2026 · 被引用 491 次
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackXingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen 等ICLR 2024 · 被引用 308 次
- Efficient Sequential Decision Making with Large Language ModelsDingyang Chen, Qi Zhang, Yinglun ZhuEMNLP 2024 · 被引用 3 次
- Promoting Topic Coherence and Inter-Document Consorts in Multi-Document Summarization via Simplicial Complex and Sheaf GraphYash Kumar Atri, Arun Iyer, Tanmoy Chakraborty, Vikram GoyalEMNLP 2023 · 被引用 2 次
相关 Paper
- From Existence to Exhaustiveness: Unveiling the Compounding Failures of LLMs in Multi-answer Event Temporal ReasoningShaojuan WuSIGIR 2026
- EvolveBench: A Comprehensive Benchmark for Assessing Temporal Awareness in LLMs on Evolving KnowledgeZhiyuan Zhu, Yusheng Liao, Zhe Chen, Yuhao Wang 等ACL 2025 · 被引用 10 次
- TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language ModelsJoel Jang, Seonghyeon Ye, Changho Lee, Sohee Yang 等EMNLP 2022 · 被引用 42 次
- ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple DomainsYein Park, Chanwoong Yoon, Jungwoo Park, Donghyeon Lee 等ICLR 2025
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in LLMsSoyeon Kim, Jindong Wang, Xing Xie, Steven Euijong WhangICLR 2026
