Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models
Qingyu Tan, Hwee Tou Ng, Lidong Bing
摘要
Reasoning about time is of fundamental importance. Many facts are time-dependent. For example, athletes change teams from time to time, and different government officials are elected periodically. Previous time-dependent question answering (QA) datasets tend to be biased in either their coverage of time spans or question types. In this paper, we introduce a comprehensive probing dataset TEMPREASON to evaluate the temporal reasoning capability of large language models. Our dataset includes questions of three temporal reasoning levels. In addition, we also propose a novel learning framework to improve the temporal reasoning capability of large language models, based on temporal span extraction and time-sensitive reinforcement learning. We conducted experiments in closed book QA, open book QA, and reasoning QA settings and demonstrated the effectiveness of our approach 1 . * Qingyu Tan is under the Joint PhD Program between Alibaba and NUS. † Corresponding author. 1 Our code and data are released on https://github.com/ DAMO-NLP-SG/TempReason
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- Large Language Models-guided Dynamic Adaptation for Temporal Knowledge Graph ReasoningJiapu Wang, Kai Sun, Linhao Luo, Wei Wei 等NeurIPS 2024 · 被引用 82 次
- Large Language Models are Temporal and Causal Reasoners for Video Question AnsweringDohwan Ko, Ji Soo Lee, Woo-Young Kang, Byungseok Roh 等EMNLP 2023 · 被引用 30 次
- Faithful Temporal Question Answering over Heterogeneous SourcesZhen Jia, Philipp Christmann, Gerhard WeikumWWW 2024 · 被引用 20 次
- HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented GenerationJie Ouyang, Tingyue Pan, Mingyue Cheng, Ruiran Yan 等ACL 2025 · 被引用 14 次
- Parenting: Optimizing Knowledge Selection of Retrieval-Augmented Language Models with Parameter Decoupling and Tailored TuningYongxin Xu, Ruizhe Zhang, Xinke Jiang, Yujie Feng 等ACL 2025 · 被引用 12 次
它引用的顶会 Paper11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- Tensor Decompositions for Temporal Knowledge Base CompletionTimothée Lacroix, Guillaume Obozinski, Nicolas UsunierICLR 2020 · 被引用 341 次
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis 等EMNLP 2020 · 被引用 142 次
相关 Paper
- Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?Zhaochen Su, Juntao Li, Jun Zhang, Tong Zhu 等ACL 2024
- Time-MQA: Time Series Multi-Task Question Answering with Context EnhancementYaxuan Kong, Yiyuan Yang, Yoontae Hwang, Wenjie Du 等ACL 2025
- TimeR⁴ : Time-aware Retrieval-Augmented Large Language Models for Temporal Knowledge Graph Question AnsweringXinying Qian, Ying Zhang, Yu Zhao, Baohang Zhou 等EMNLP 2024 · 被引用 11 次
- Learning to Reason Over Time: Timeline Self-Reflection for Improved Temporal Reasoning in Language ModelsAdrián Bazaga, Rexhina Blloshmi, Bill Byrne, Adrià de GispertACL 2025
- Temporal Knowledge Question Answering via Abstract Reasoning InductionZiyang Chen, Dongfang Li, Xiang Zhao, Baotian Hu 等ACL 2024 · 被引用 9 次
