Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models
Qingyu Tan, Hwee Tou Ng, Lidong Bing
Abstract
Reasoning about time is of fundamental importance. Many facts are time-dependent. For example, athletes change teams from time to time, and different government officials are elected periodically. Previous time-dependent question answering (QA) datasets tend to be biased in either their coverage of time spans or question types. In this paper, we introduce a comprehensive probing dataset TEMPREASON to evaluate the temporal reasoning capability of large language models. Our dataset includes questions of three temporal reasoning levels. In addition, we also propose a novel learning framework to improve the temporal reasoning capability of large language models, based on temporal span extraction and time-sensitive reinforcement learning. We conducted experiments in closed book QA, open book QA, and reasoning QA settings and demonstrated the effectiveness of our approach 1 . * Qingyu Tan is under the Joint PhD Program between Alibaba and NUS. † Corresponding author. 1 Our code and data are released on https://github.com/ DAMO-NLP-SG/TempReason
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cef3fbf0-fdd2-4228-b3c5-66a0e49bcc27Cited by top-tier papers35
- Large Language Models-guided Dynamic Adaptation for Temporal Knowledge Graph ReasoningJiapu Wang, Kai Sun, Linhao Luo, Wei Wei et al.NeurIPS 2024 · 82 citations
- Large Language Models are Temporal and Causal Reasoners for Video Question AnsweringDohwan Ko, Ji Soo Lee, Woo-Young Kang, Byungseok Roh et al.EMNLP 2023 · 30 citations
- Faithful Temporal Question Answering over Heterogeneous SourcesZhen Jia, Philipp Christmann, Gerhard WeikumWWW 2024 · 20 citations
- HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented GenerationJie Ouyang, Tingyue Pan, Mingyue Cheng, Ruiran Yan et al.ACL 2025 · 14 citations
- Parenting: Optimizing Knowledge Selection of Retrieval-Augmented Language Models with Parameter Decoupling and Tailored TuningYongxin Xu, Ruizhe Zhang, Xinke Jiang, Yujie Feng et al.ACL 2025 · 12 citations
Builds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Tensor Decompositions for Temporal Knowledge Base CompletionTimothée Lacroix, Guillaume Obozinski, Nicolas UsunierICLR 2020 · 341 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?Zhaochen Su, Juntao Li, Jun Zhang, Tong Zhu et al.ACL 2024
- Time-MQA: Time Series Multi-Task Question Answering with Context EnhancementYaxuan Kong, Yiyuan Yang, Yoontae Hwang, Wenjie Du et al.ACL 2025
- TimeR⁴ : Time-aware Retrieval-Augmented Large Language Models for Temporal Knowledge Graph Question AnsweringXinying Qian, Ying Zhang, Yu Zhao, Baohang Zhou et al.EMNLP 2024 · 11 citations
- Learning to Reason Over Time: Timeline Self-Reflection for Improved Temporal Reasoning in Language ModelsAdrián Bazaga, Rexhina Blloshmi, Bill Byrne, Adrià de GispertACL 2025
- Temporal Knowledge Question Answering via Abstract Reasoning InductionZiyang Chen, Dongfang Li, Xiang Zhao, Baotian Hu et al.ACL 2024 · 9 citations
