SPAN: Benchmarking and Improving Cross-Calendar Temporal Reasoning of Large Language Models
Zhongjian Miao, Hao Fu, Chen Wei
Abstract
Temporal reasoning is a fundamental capability for large language models (LLMs) to understand real-world dynamics. Existing research on temporal reasoning has predominantly focused on the Gregorian calendar. However, as many countries and regions concurrently adopt multiple calendar systems, temporal reasoning across calendars becomes crucial for LLMs in global and multicultural contexts. Unfortunately, cross-calendar temporal reasoning remains underexplored, with no dedicated benchmark available to evaluate this capability. To bridge this gap, we introduce SPAN, a croSs-calendar temPoral reAsoning beNchmark, which requires LLMs to perform intra-calendar temporal reasoning and inter-calendar temporal conversion. SPAN features ten cross-calendar temporal reasoning directions, two reasoning types, and two question formats across six calendars. To enable time-variant and contamination-free evaluation, we propose a template-driven protocol for dynamic instance generation that enables assessment on a user-specified Gregorian date. We conduct extensive experiments on both open-and closed-source state-of-theart (SOTA) LLMs over a range of dates spanning 100 years from 1960 to 2060. Our evaluations show that these LLMs achieve an average accuracy of only 34.5%, with none exceeding 80%, indicating that this task remains challenging. Through in-depth analysis of reasoning types, question formats, and temporal reasoning directions, we identify two key obstacles for LLMs: Future-Date Degradation and Calendar Asymmetry Bias. To strengthen LLMs' cross-calendar temporal reasoning capability, we further develop an LLM-powered Time Agent that leverages tool-augmented code generation. Empirical results show that Time Agent achieves an average accuracy of 95.31%, outperforming several competitive baselines, highlighting the potential of tool-augmented code generation to advance cross-calendar temporal reasoning. We hope this work will inspire further efforts toward more temporally and culturally adaptive LLMs. Our source code and datasets are available at https://github.com/miaozhongjian/span.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- StreamingQA: A Benchmark for Adaptation to New Knowledge over Time in Question Answering ModelsAdam Liska, Tomás Kociský, Elena Gribovskaya, Tayfun Terzi et al.ICML 2022 · 129 citations
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha et al.EMNLP 2023 · 17 citations
- TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language ModelsZheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu et al.ACL 2024 · 12 citations
- Test of Time: A Benchmark for Evaluating LLMs on Temporal ReasoningBahare Fatemi, Mehran Kazemi, Anton Tsitsulin, Karishma Malkan et al.ICLR 2025 · 2 citations
Related papers
- TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at ScaleMalgorzata Gwiazda, Yifu Cai, Mononito Goswami, Arjun Choudhry et al.ICLR 2026 · 6 citations
- On Path to Multimodal Historical Reasoning: HistBench and HistAgentJiahao Qiu, Fulian Xiao, Yimin Wang, Yuchen Mao et al.ICML 2026 · 5 citations
- Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language ModelsQingyu Tan, Hwee Tou Ng, Lidong BingACL 2023 · 24 citations
- Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?Zhaochen Su, Juntao Li, Jun Zhang, Tong Zhu et al.ACL 2024
- TCP: a Benchmark for Temporal Constraint-Based PlanningZifeng Ding, Sikuan Yan, Moy Yuan, Xianglong Hu et al.EMNLP 2025
