TIMEDIAL: Temporal Commonsense Reasoning in Dialog
Lianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He, Yejin Choi, Manaal Faruqui
Abstract
Everyday conversations require understanding everyday events, which in turn, requires understanding temporal commonsense concepts interwoven with those events. Despite recent progress with massive pre-trained language models (LMs) such as T5 and GPT-3, their capability of temporal reasoning in dialogs remains largely under-explored. In this paper, we present the first study to investigate pre-trained LMs for their temporal reasoning capabilities in dialogs by introducing a new task and a crowd-sourced English challenge set, TIMEDIAL. We formulate TIME-DIAL as a multiple choice cloze task with over 1.1K carefully curated dialogs. Empirical results demonstrate that even the best performing models struggle on this task compared to humans, with 23 absolute points of gap in accuracy. Furthermore, our analysis reveals that the models fail to reason about dialog context correctly; instead, they rely on shallow cues based on existing temporal patterns in context, motivating future research for modeling temporal concepts in text and robust contextual reasoning about them. The dataset is publicly available at: https://github.com/ google-research-datasets/timedial .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f7f9ddf-cba7-4b65-ba37-339ba83d3787Cited by top-tier papers23
- InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction TuningPrakhar Gupta, Cathy Jiao, Yi-Ting Yeh, Shikib Mehri et al.EMNLP 2022 · 26 citations
- Knowledge Infused DecodingRuibo Liu, Guoqing Zheng, Shashank Gupta, Radhika Gaonkar et al.ICLR 2022 · 18 citations
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha et al.EMNLP 2023 · 17 citations
- Reverse Multi-Choice Dialogue Commonsense Inference with Graph-of-ThoughtLi Zheng, Hao Fei, Fei Li, Bobo Li et al.AAAI 2024 · 13 citations
- Mitigating Reporting Bias in Semi-supervised Temporal Commonsense Inference with Probabilistic Soft LogicBibo Cai, Xiao Ding, Bowen Chen, Li Du et al.AAAI 2022 · 12 citations
Builds on6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers et al.ICML 2020 · 242 citations
- Inducing Relational Knowledge from BERTZied Bouraoui, José Camacho-Collados, Steven SchockaertAAAI 2020 · 183 citations
- TORQUE: A Reading Comprehension Dataset of Temporal Ordering QuestionsQiang Ning, Hao Wu, Rujun Han, Nanyun Peng et al.EMNLP 2020 · 79 citations
- Temporal Common Sense Acquisition with Minimal SupervisionBen Zhou, Qiang Ning, Daniel Khashabi, Dan RothACL 2020 · 76 citations
Related papers
- Self-Supervised Logic Induction for Explainable Fuzzy Temporal Commonsense ReasoningBibo Cai, Xiao Ding, Zhouhao Sun, Bing Qin et al.AAAI 2023 · 11 citations
- Around the World in 24 Hours: Probing LLM Knowledge of Time and PlaceCarolin Holtermann, Paul Röttger, Anne LauscherACL 2025
- TGEA: An Error-Annotated Dataset and Benchmark Tasks for TextGeneration from Pretrained Language ModelsJie He, Bo Peng, Yi Liao, Qun Liu et al.ACL 2021
- Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?Zhaochen Su, Juntao Li, Jun Zhang, Tong Zhu et al.ACL 2024
- TCP: a Benchmark for Temporal Constraint-Based PlanningZifeng Ding, Sikuan Yan, Moy Yuan, Xianglong Hu et al.EMNLP 2025
