Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?
Zhaochen Su, Juntao Li, Jun Zhang, Tong Zhu, Xiaoye Qu, Pan Zhou, Yan Bowen, Yu Cheng, Min Zhang
Abstract
Temporal reasoning is fundamental for large language models (LLMs) to comprehend the world. Current temporal reasoning datasets are limited to questions about single or isolated events, falling short in mirroring the realistic temporal characteristics involving concurrent nature and intricate temporal interconnections. In this paper, we introduce COTEMP-QA, a comprehensive co-temporal Question Answering (QA) benchmark containing four co-temporal scenarios (Equal, Overlap, During, Mix) with 4,749 samples for evaluating the co-temporal comprehension and reasoning abilities of LLMs. Our extensive experiments reveal a significant gap between the performance of current LLMs and human-level reasoning on COTEMPQA tasks. Even when enhanced with Chain of Thought (CoT) methodologies, models consistently struggle with our task. In our preliminary exploration, we discovered that mathematical reasoning plays a significant role in handling co-temporal events and proposed a strategy to boost LLMs' co-temporal reasoning from a mathematical perspective. We hope that our COTEMPQA datasets will encourage further advancements in improving the co-temporal reasoning capabilities of LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af4e03b7-bc23-4a85-a1d5-b2b18da6a3a6Cited by top-tier papers4
- EvolveBench: A Comprehensive Benchmark for Assessing Temporal Awareness in LLMs on Evolving KnowledgeZhiyuan Zhu, Yusheng Liao, Zhe Chen, Yuhao Wang et al.ACL 2025 · 10 citations
- It's High Time: A Survey of Temporal Question AnsweringBhawna Piryani, Abdelrahman Abdallah, Jamshid Mozafari, Avishek Anand et al.ACL 2026 · 6 citations
- PIPER: Benchmarking and Prompting Event Reasoning Boundary of LLMs via Debiasing-Distillation Enhanced TuningZhicong Lu, Changyuan Tian, PeiguangLi PeiguangLi, Li Jin et al.ACL 2025 · 4 citations
- From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data CalibrationMingyang Song, Xiaoye Qu, Jiawei Zhou, Yu ChengCVPR 2025
Builds on7
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- StreamingQA: A Benchmark for Adaptation to New Knowledge over Time in Question Answering ModelsAdam Liska, Tomás Kociský, Elena Gribovskaya, Tayfun Terzi et al.ICML 2022 · 129 citations
- ECONET: Effective Continual Pretraining of Language Models for Event Temporal ReasoningRujun Han, Xiang Ren, Nanyun PengEMNLP 2021 · 31 citations
- Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language ModelsQingyu Tan, Hwee Tou Ng, Lidong BingACL 2023 · 24 citations
- Solving Math Word Problems via Cooperative Reasoning induced Language ModelsXinyu Zhu, Junjie Wang, Lin Zhang, Yuxiang Zhang et al.ACL 2023 · 16 citations
Related papers
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha et al.EMNLP 2023 · 17 citations
- UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' MemorizationMd Nayem Uddin, Amir Saeidi, Divij Handa, Agastya Seth et al.ACL 2025 · 18 citations
- ComplexTempQA: A 100m Dataset for Complex Temporal Question AnsweringRaphael Gruber, Abdelrahman Abdallah, Michael Färber, Adam JatowtEMNLP 2025 · 2 citations
- Test of Time: A Benchmark for Evaluating LLMs on Temporal ReasoningBahare Fatemi, Mehran Kazemi, Anton Tsitsulin, Karishma Malkan et al.ICLR 2025 · 2 citations
- TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language ModelsZheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu et al.ACL 2024 · 12 citations
