Learning to Reason Over Time: Timeline Self-Reflection for Improved Temporal Reasoning in Language Models
Adrián Bazaga, Rexhina Blloshmi, Bill Byrne, Adrià de Gispert
Abstract
Large Language Models (LLMs) have emerged as powerful tools for generating coherent text, understanding context, and performing reasoning tasks. However, they struggle with temporal reasoning, which requires processing time-related information such as event sequencing, durations, and inter-temporal relationships. These capabilities are critical for applications including question answering, scheduling, and historical analysis. In this paper, we introduce TISER, a novel framework that enhances the temporal reasoning abilities of LLMs through a multi-stage process that combines timeline construction with iterative self-reflection. Our approach leverages test-time scaling to extend the length of reasoning traces, enabling models to capture complex temporal dependencies more effectively. This strategy not only boosts reasoning accuracy but also improves the traceability of the inference process. Experimental results demonstrate state-of-the-art performance across multiple benchmarks, including out-of-distribution test sets, and reveal that TISER enables smaller open-source models to surpass larger closed-weight models on challenging temporal reasoning tasks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9dc9e92-9839-46ed-a27d-345aa4a6ef04Cited by top-tier papers4
- REMem: Reasoning with Episodic Memory in Language AgentYiheng Shu, Padmaja Jonnalagedda, Xiang Gao, Bernal Jimenez Gutierrez et al.ICLR 2026 · 20 citations
- It's High Time: A Survey of Temporal Question AnsweringBhawna Piryani, Abdelrahman Abdallah, Jamshid Mozafari, Avishek Anand et al.ACL 2026 · 6 citations
- ODL-TempLLM: Ontology-Guided and Description Logic-Reasoned Temporal Reasoning with LLMsJinshuo Liu, Cheng Bi, Meng Wang, Juan Deng et al.ACL 2026
- NeSTR: A Neuro-Symbolic Abductive Framework for Temporal Reasoning in Large Language ModelsFeng Liang, Weixin Zeng, Runhao Zhao, Xiang ZhaoAAAI 2026
Builds on10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- Large Language Models Can Learn Temporal ReasoningSiheng Xiong, Ali Payani, Ramana Kompella, Faramarz FekriACL 2024
- TimelineReasoner: Advancing Timeline Summarization with Large Reasoning ModelsLiancheng Zhang, Xiaoxi Li, Zhicheng DouSIGIR 2026
- Temporal Knowledge Question Answering via Abstract Reasoning InductionZiyang Chen, Dongfang Li, Xiang Zhao, Baotian Hu et al.ACL 2024 · 9 citations
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha et al.EMNLP 2023 · 17 citations
- Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language ModelsQingyu Tan, Hwee Tou Ng, Lidong BingACL 2023 · 24 citations
