TORQUE: A Reading Comprehension Dataset of Temporal Ordering Questions
Qiang Ning, Hao Wu, Rujun Han, Nanyun Peng, Matt Gardner, Dan Roth
Abstract
A critical part of reading is being able to understand the temporal relationships between events described in a passage of text, even when those relationships are not explicitly stated. However, current machine reading comprehension benchmarks have practically no questions that test temporal phenomena, so systems trained on these benchmarks have no capacity to answer questions such as "what happened before/after [some event]?" We introduce TORQUE, a new English reading comprehension benchmark built on 3.2k news snippets with 21k human-generated questions querying temporal relationships. Results show that RoBERTa-large achieves an exact-match score of 51% on the test set of TORQUE, about 30% behind human performance. 1 1 https://allennlp.org/torque.html Heavy snow is causing disruption to transport across the UK, with heavy rainfall bringing flooding to the south-west of England. Rescuers searching for a woman trapped in a landslide at her home in Looe, Cornwall, said they had found a body. Q1: What events have already finished? A: searching trapped landslide said found Q2: What events have begun but has not finished? A: snow causing disruption rainfall bringing flooding Q3: What will happen in the future? A: No answers. Q4: What happened before a woman was trapped? A: landslide Q5: What had started before a woman was trapped? A: snow rainfall landslide Q6: What happened while a woman was trapped? A: searching Q7: What happened after a woman was trapped? A: searching said found Q8: What happened at about the same time as the snow? A: rainfall Q9: What happened after the snow started? A: causing disruption bringing flooding searching trapped landslide said found Q10: What happened before the snow started? A: No answers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers27
- StreamingQA: A Benchmark for Adaptation to New Knowledge over Time in Question Answering ModelsAdam Liska, Tomás Kociský, Elena Gribovskaya, Tayfun Terzi et al.ICML 2022 · 129 citations
- Language Models Can Improve Event Prediction by Few-Shot Abductive ReasoningXiaoming Shi, Siqiao Xue, Kangrui Wang, Fan Zhou et al.NeurIPS 2023 · 95 citations
- Improving Time Sensitivity for Question Answering over Temporal Knowledge GraphsChao Shang, Guangtao Wang, Peng Qi, Jing HuangACL 2022 · 55 citations
- Faking Fake News for Real Fake News Detection: Propaganda-Loaded Training Data GenerationKung-Hsiang Huang, Kathleen R. McKeown, Preslav Nakov, Yejin Choi et al.ACL 2023 · 35 citations
- ECONET: Effective Continual Pretraining of Language Models for Event Temporal ReasoningRujun Han, Xiang Ren, Nanyun PengEMNLP 2021 · 31 citations
Builds on2
Related papers
- Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?Zhaochen Su, Juntao Li, Jun Zhang, Tong Zhu et al.ACL 2024
- Will LLMs Replace the Encoder-Only Models in Temporal Relation Classification?Gabriel Roccabruna, Massimo Rizzoli, Giuseppe RiccardiEMNLP 2024 · 3 citations
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha et al.EMNLP 2023 · 17 citations
- ESTER: A Machine Reading Comprehension Dataset for Reasoning about Event Semantic RelationsRujun Han, I-Hung Hsu, Jiao Sun, Julia Baylon et al.EMNLP 2021 · 30 citations
- Analyzing Temporal Complex Events with Large Language Models? A Benchmark towards Temporal, Long Context UnderstandingZhihan Zhang, Yixin Cao, Chenchen Ye, Yunshan Ma et al.ACL 2024
