Generic Temporal Reasoning with Differential Analysis and Explanation
Yu Feng, Ben Zhou, Haoyu Wang, Helen Jin, Dan Roth
Abstract
Temporal reasoning is the task of predicting temporal relations of event pairs. While temporal reasoning models can perform reasonably well on in-domain benchmarks, we have little idea of these systems’ generalizability due to existing datasets’ limitations. In this work, we introduce a novel task named TODAY that bridges this gap with temporal differential analysis, which as the name suggests, evaluates whether systems can correctly understand the effect of incremental changes. Specifically, TODAY introduces slight contextual changes for given event pairs, and systems are asked to tell how this subtle contextual change would affect relevant temporal relation distributions. To facilitate learning, TODAY also annotates human explanations. We show that existing models, including GPT-3.5, drop to random guessing on TODAY, suggesting that they heavily rely on spurious information rather than proper reasoning for temporal predictions. On the other hand, we show that TODAY’s supervision style and explanation annotations can be used in joint learning, encouraging models to use more appropriate signals during training and thus outperform across several benchmarks. TODAY can also be used to train models to solicit incidental supervision from noisy sources such as GPT-3.5, thus moving us more toward the goal of generic temporal reasoning systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Language Models Can Improve Event Prediction by Few-Shot Abductive ReasoningXiaoming Shi, Siqiao Xue, Kangrui Wang, Fan Zhou et al.NeurIPS 2023 · 95 citations
- Complex Query Answering on Eventuality Knowledge Graph with Implicit Logical ConstraintsJiaxin Bai, Xin Liu, Weiqi Wang, Chen Luo et al.NeurIPS 2023 · 46 citations
- It's High Time: A Survey of Temporal Question AnsweringBhawna Piryani, Abdelrahman Abdallah, Jamshid Mozafari, Avishek Anand et al.ACL 2026 · 6 citations
- Temporal reasoning for timeline summarisation in social mediaJiayu Song, Mahmud Elahi Akhter, Dana Atzil-Slonim, Maria LiakataACL 2025 · 6 citations
- Every Answer Matters: Evaluating Commonsense with Probabilistic MeasuresQi Cheng, Michael Boratko, Pranay Kumar Yelugam, Tim O'Gorman et al.ACL 2024 · 2 citations
Builds on9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- TORQUE: A Reading Comprehension Dataset of Temporal Ordering QuestionsQiang Ning, Hao Wu, Rujun Han, Nanyun Peng et al.EMNLP 2020 · 79 citations
- Temporal Common Sense Acquisition with Minimal SupervisionBen Zhou, Qiang Ning, Daniel Khashabi, Dan RothACL 2020 · 76 citations
- Selecting Optimal Context Sentences for Event-Event Relation ExtractionHieu Man, Nghia Trung Ngo, Linh Ngo Van, Thien Huu NguyenAAAI 2022 · 59 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
Related papers
- TIMEDIAL: Temporal Commonsense Reasoning in DialogLianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He et al.ACL 2021
- Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?Zhaochen Su, Juntao Li, Jun Zhang, Tong Zhu et al.ACL 2024
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha et al.EMNLP 2023 · 17 citations
- Revisiting Relation Extraction in the era of Large Language ModelsSomin Wadhwa, Silvio Amir, Byron C. WallaceACL 2023 · 145 citations
- The Unreliability of Explanations in Few-shot Prompting for Textual ReasoningXi Ye, Greg DurrettNeurIPS 2022 · 272 citations
