Back to the Future: Towards Explainable Temporal Reasoning with Large Language Models
Chenhan Yuan, Qianqian Xie, Jimin Huang, Sophia Ananiadou
Abstract
Temporal reasoning is a crucial natural language processing (NLP) task, providing a nuanced understanding of time-sensitive contexts within textual data. Although recent advancements in Large Language Models (LLMs) have demonstrated their potential in temporal reasoning, the predominant focus has been on tasks such as temporal expression detection, normalization, and temporal relation extraction. These tasks are primarily designed for the extraction of direct and past temporal cues from given contexts and to engage in simple reasoning processes. A significant gap remains when considering complex reasoning tasks such as event forecasting, which requires multi-step temporal reasoning on events and prediction on the future timestamp. Another notable limitation of existing methods is their incapability to provide an illustration of their reasoning process for explaining their prediction, hindering explainability. In this paper, we introduce the first task of explainable temporal reasoning, to predict an event's occurrence at a future timestamp based on context which requires multiple reasoning over multiple events, and subsequently provide a clear explanation for their prediction. Our task offers a comprehensive evaluation of both the LLMs' complex temporal reasoning ability, the future event prediction ability, and explainability-a critical attribute for AI applications. To support this task, we present the first multi-source instruction-tuning dataset of explainable temporal reasoning (ExpTime) with 26k derived from the temporal knowledge graph datasets and their temporal reasoning paths, using a novel knowledge-graph-instructedgeneration strategy. Based on the dataset, we propose the first open-source LLM series TimeLlaMA based on the foundation LLM LlaMA2, with the ability of instruction following for explainable temporal reasoning. We compare the performance of our method and a variety of LLMs, where our method achieves the state-of-theart performance of temporal prediction and explanation generation. We also explore the impact of instruction tuning and different training sizes of instruction-tuning data, highlighting LLM's capabilities and limitations in complex temporal prediction and explanation generation. We release our datasets, models, and evaluations to facilitate future research at https://github.com/chenhan97/TimeLlama . CCS CONCEPTS • Computing methodologies → Temporal reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 593e937e-d37d-4e2e-a962-a80105c371bbCited by top-tier papers13
- Large Language Models-guided Dynamic Adaptation for Temporal Knowledge Graph ReasoningJiapu Wang, Kai Sun, Linhao Luo, Wei Wei et al.NeurIPS 2024 · 82 citations
- AssoMem: Scalable Memory QA with Multi-Signal Associative RetrievalKai Zhang, Xinyuan Zhang, Ejaz Ahmed, Hongda Jiang et al.ICLR 2026 · 9 citations
- Beyond Single Pass, Looping Through Time: KG-IRAG with Iterative Knowledge RetrievalRuiyi Yang, Hao Xue, Imran Razzak, Flora D. SalimWWW 2026 · 8 citations
- Improving Large Language Models in Event Relation Logical PredictionMeiqi Chen, Yubo Ma, Kaitao Song, Yixin Cao et al.ACL 2024 · 7 citations
- CoMaPOI: A Collaborative Multi-Agent Framework for Next POI Prediction Bridging the Gap Between Trajectory and LanguageLin Zhong, Lingzhi Wang, Xu Yang, Qing LiaoSIGIR 2025 · 6 citations
Builds on22
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Recurrent Event Network: Autoregressive Structure Inferenceover Temporal Knowledge GraphsWoojeong Jin, Meng Qu, Xisen Jin, Xiang RenEMNLP 2020 · 353 citations
Related papers
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha et al.EMNLP 2023 · 17 citations
- Will LLMs Replace the Encoder-Only Models in Temporal Relation Classification?Gabriel Roccabruna, Massimo Rizzoli, Giuseppe RiccardiEMNLP 2024 · 3 citations
- ODL-TempLLM: Ontology-Guided and Description Logic-Reasoned Temporal Reasoning with LLMsJinshuo Liu, Cheng Bi, Meng Wang, Juan Deng et al.ACL 2026
- Fostering Video Reasoning via Next-Event PredictionHaonan Wang, Hongfu Liu, Xiangyan Liu, Chao Du et al.ICLR 2026 · 14 citations
- TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language ModelsZheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu et al.ACL 2024 · 12 citations
