Generic Temporal Reasoning with Differential Analysis and Explanation
Yu Feng, Ben Zhou, Haoyu Wang, Helen Jin, Dan Roth
摘要
Temporal reasoning is the task of predicting temporal relations of event pairs. While temporal reasoning models can perform reasonably well on in-domain benchmarks, we have little idea of these systems’ generalizability due to existing datasets’ limitations. In this work, we introduce a novel task named TODAY that bridges this gap with temporal differential analysis, which as the name suggests, evaluates whether systems can correctly understand the effect of incremental changes. Specifically, TODAY introduces slight contextual changes for given event pairs, and systems are asked to tell how this subtle contextual change would affect relevant temporal relation distributions. To facilitate learning, TODAY also annotates human explanations. We show that existing models, including GPT-3.5, drop to random guessing on TODAY, suggesting that they heavily rely on spurious information rather than proper reasoning for temporal predictions. On the other hand, we show that TODAY’s supervision style and explanation annotations can be used in joint learning, encouraging models to use more appropriate signals during training and thus outperform across several benchmarks. TODAY can also be used to train models to solicit incidental supervision from noisy sources such as GPT-3.5, thus moving us more toward the goal of generic temporal reasoning systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Language Models Can Improve Event Prediction by Few-Shot Abductive ReasoningXiaoming Shi, Siqiao Xue, Kangrui Wang, Fan Zhou 等NeurIPS 2023 · 被引用 95 次
- Complex Query Answering on Eventuality Knowledge Graph with Implicit Logical ConstraintsJiaxin Bai, Xin Liu, Weiqi Wang, Chen Luo 等NeurIPS 2023 · 被引用 46 次
- It's High Time: A Survey of Temporal Question AnsweringBhawna Piryani, Abdelrahman Abdallah, Jamshid Mozafari, Avishek Anand 等ACL 2026 · 被引用 6 次
- Temporal reasoning for timeline summarisation in social mediaJiayu Song, Mahmud Elahi Akhter, Dana Atzil-Slonim, Maria LiakataACL 2025 · 被引用 6 次
- Every Answer Matters: Evaluating Commonsense with Probabilistic MeasuresQi Cheng, Michael Boratko, Pranay Kumar Yelugam, Tim O'Gorman 等ACL 2024 · 被引用 2 次
它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- TORQUE: A Reading Comprehension Dataset of Temporal Ordering QuestionsQiang Ning, Hao Wu, Rujun Han, Nanyun Peng 等EMNLP 2020 · 被引用 79 次
- Temporal Common Sense Acquisition with Minimal SupervisionBen Zhou, Qiang Ning, Daniel Khashabi, Dan RothACL 2020 · 被引用 76 次
- Selecting Optimal Context Sentences for Event-Event Relation ExtractionHieu Man, Nghia Trung Ngo, Linh Ngo Van, Thien Huu NguyenAAAI 2022 · 被引用 59 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
相关 Paper
- TIMEDIAL: Temporal Commonsense Reasoning in DialogLianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He 等ACL 2021
- Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?Zhaochen Su, Juntao Li, Jun Zhang, Tong Zhu 等ACL 2024
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha 等EMNLP 2023 · 被引用 17 次
- Revisiting Relation Extraction in the era of Large Language ModelsSomin Wadhwa, Silvio Amir, Byron C. WallaceACL 2023 · 被引用 145 次
- The Unreliability of Explanations in Few-shot Prompting for Textual ReasoningXi Ye, Greg DurrettNeurIPS 2022 · 被引用 272 次
