Back to the Future: Towards Explainable Temporal Reasoning with Large Language Models
Chenhan Yuan, Qianqian Xie, Jimin Huang, Sophia Ananiadou
摘要
Temporal reasoning is a crucial natural language processing (NLP) task, providing a nuanced understanding of time-sensitive contexts within textual data. Although recent advancements in Large Language Models (LLMs) have demonstrated their potential in temporal reasoning, the predominant focus has been on tasks such as temporal expression detection, normalization, and temporal relation extraction. These tasks are primarily designed for the extraction of direct and past temporal cues from given contexts and to engage in simple reasoning processes. A significant gap remains when considering complex reasoning tasks such as event forecasting, which requires multi-step temporal reasoning on events and prediction on the future timestamp. Another notable limitation of existing methods is their incapability to provide an illustration of their reasoning process for explaining their prediction, hindering explainability. In this paper, we introduce the first task of explainable temporal reasoning, to predict an event's occurrence at a future timestamp based on context which requires multiple reasoning over multiple events, and subsequently provide a clear explanation for their prediction. Our task offers a comprehensive evaluation of both the LLMs' complex temporal reasoning ability, the future event prediction ability, and explainability-a critical attribute for AI applications. To support this task, we present the first multi-source instruction-tuning dataset of explainable temporal reasoning (ExpTime) with 26k derived from the temporal knowledge graph datasets and their temporal reasoning paths, using a novel knowledge-graph-instructedgeneration strategy. Based on the dataset, we propose the first open-source LLM series TimeLlaMA based on the foundation LLM LlaMA2, with the ability of instruction following for explainable temporal reasoning. We compare the performance of our method and a variety of LLMs, where our method achieves the state-of-theart performance of temporal prediction and explanation generation. We also explore the impact of instruction tuning and different training sizes of instruction-tuning data, highlighting LLM's capabilities and limitations in complex temporal prediction and explanation generation. We release our datasets, models, and evaluations to facilitate future research at https://github.com/chenhan97/TimeLlama . CCS CONCEPTS • Computing methodologies → Temporal reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Large Language Models-guided Dynamic Adaptation for Temporal Knowledge Graph ReasoningJiapu Wang, Kai Sun, Linhao Luo, Wei Wei 等NeurIPS 2024 · 被引用 82 次
- AssoMem: Scalable Memory QA with Multi-Signal Associative RetrievalKai Zhang, Xinyuan Zhang, Ejaz Ahmed, Hongda Jiang 等ICLR 2026 · 被引用 9 次
- Beyond Single Pass, Looping Through Time: KG-IRAG with Iterative Knowledge RetrievalRuiyi Yang, Hao Xue, Imran Razzak, Flora D. SalimWWW 2026 · 被引用 8 次
- Improving Large Language Models in Event Relation Logical PredictionMeiqi Chen, Yubo Ma, Kaitao Song, Yixin Cao 等ACL 2024 · 被引用 7 次
- CoMaPOI: A Collaborative Multi-Agent Framework for Next POI Prediction Bridging the Gap Between Trajectory and LanguageLin Zhong, Lingzhi Wang, Xu Yang, Qing LiaoSIGIR 2025 · 被引用 6 次
它引用的顶会 Paper22
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Recurrent Event Network: Autoregressive Structure Inferenceover Temporal Knowledge GraphsWoojeong Jin, Meng Qu, Xisen Jin, Xiang RenEMNLP 2020 · 被引用 353 次
相关 Paper
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha 等EMNLP 2023 · 被引用 17 次
- Will LLMs Replace the Encoder-Only Models in Temporal Relation Classification?Gabriel Roccabruna, Massimo Rizzoli, Giuseppe RiccardiEMNLP 2024 · 被引用 3 次
- ODL-TempLLM: Ontology-Guided and Description Logic-Reasoned Temporal Reasoning with LLMsJinshuo Liu, Cheng Bi, Meng Wang, Juan Deng 等ACL 2026
- Fostering Video Reasoning via Next-Event PredictionHaonan Wang, Hongfu Liu, Xiangyan Liu, Chao Du 等ICLR 2026 · 被引用 14 次
- TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language ModelsZheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu 等ACL 2024 · 被引用 12 次
