UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
Md Nayem Uddin, Amir Saeidi, Divij Handa, Agastya Seth, Tran Cao Son, Eduardo Blanco, Steven R. Corman, Chitta Baral
Abstract
This paper introduces UnSeenTimeQA, a novel data contamination-free time-sensitive question-answering (TSQA) benchmark. It differs from existing TSQA benchmarks by avoiding web-searchable queries grounded in the real world. We present a series of timesensitive event scenarios based on synthetically generated facts. It requires large language models (LLMs) to engage in genuine temporal reasoning without depending on the factual knowledge acquired during the pre-training phase. Our data generation framework enables ondemand generation of new samples, mitigating the risk of data leakage. We designed three types of time-sensitive questions to test LLMs' temporal reasoning abilities over sequential and parallel event occurrences. Our evaluation of five LLMs on synthetic fact-based TSQA reveals mixed results: while they perform well on simpler subsets, their overall performance remains inferior as compared to real world fact-based TSQA. Error analysis indicates that LLMs face difficulties in reasoning over longrange event dependencies and parallel events. TimeQA Temp Reason MenatQA UnSeen TimeQA (Ours)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c650e9dc-a3d6-4830-88e2-decee9e0d54fCited by top-tier papers5
- It's High Time: A Survey of Temporal Question AnsweringBhawna Piryani, Abdelrahman Abdallah, Jamshid Mozafari, Avishek Anand et al.ACL 2026 · 6 citations
- Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic EvaluationSimin Chen, Yiming Chen, Zexin Li, Yifan Jiang et al.EMNLP 2025 · 2 citations
- Beyond Timestamps: Bridging Forward and Backward Reasoning in Temporal Numerical and Relational UnderstandingXinying Qian, Ying Zhang, Xuhui Sui, Yu Zhao et al.ACL 2026
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in LLMsSoyeon Kim, Jindong Wang, Xing Xie, Steven Euijong WhangICLR 2026
- ActionReasoningBench: Reasoning about Actions with and without Ramification ConstraintsDivij Handa, Pavel Dolin, Shrinidhi Kumbhar, Tran Cao Son et al.ICLR 2025
Related papers
- Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?Zhaochen Su, Juntao Li, Jun Zhang, Tong Zhu et al.ACL 2024
- Time-MQA: Time Series Multi-Task Question Answering with Context EnhancementYaxuan Kong, Yiyuan Yang, Yoontae Hwang, Wenjie Du et al.ACL 2025
- Test of Time: A Benchmark for Evaluating LLMs on Temporal ReasoningBahare Fatemi, Mehran Kazemi, Anton Tsitsulin, Karishma Malkan et al.ICLR 2025 · 2 citations
- TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at ScaleMalgorzata Gwiazda, Yifu Cai, Mononito Goswami, Arjun Choudhry et al.ICLR 2026 · 6 citations
- CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmarkJian Wu, Linyi Yang, Zhen Wang, Manabu Okumura et al.ICLR 2025
