TORQUE: A Reading Comprehension Dataset of Temporal Ordering Questions
Qiang Ning, Hao Wu, Rujun Han, Nanyun Peng, Matt Gardner, Dan Roth
摘要
A critical part of reading is being able to understand the temporal relationships between events described in a passage of text, even when those relationships are not explicitly stated. However, current machine reading comprehension benchmarks have practically no questions that test temporal phenomena, so systems trained on these benchmarks have no capacity to answer questions such as "what happened before/after [some event]?" We introduce TORQUE, a new English reading comprehension benchmark built on 3.2k news snippets with 21k human-generated questions querying temporal relationships. Results show that RoBERTa-large achieves an exact-match score of 51% on the test set of TORQUE, about 30% behind human performance. 1 1 https://allennlp.org/torque.html Heavy snow is causing disruption to transport across the UK, with heavy rainfall bringing flooding to the south-west of England. Rescuers searching for a woman trapped in a landslide at her home in Looe, Cornwall, said they had found a body. Q1: What events have already finished? A: searching trapped landslide said found Q2: What events have begun but has not finished? A: snow causing disruption rainfall bringing flooding Q3: What will happen in the future? A: No answers. Q4: What happened before a woman was trapped? A: landslide Q5: What had started before a woman was trapped? A: snow rainfall landslide Q6: What happened while a woman was trapped? A: searching Q7: What happened after a woman was trapped? A: searching said found Q8: What happened at about the same time as the snow? A: rainfall Q9: What happened after the snow started? A: causing disruption bringing flooding searching trapped landslide said found Q10: What happened before the snow started? A: No answers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- StreamingQA: A Benchmark for Adaptation to New Knowledge over Time in Question Answering ModelsAdam Liska, Tomás Kociský, Elena Gribovskaya, Tayfun Terzi 等ICML 2022 · 被引用 129 次
- Language Models Can Improve Event Prediction by Few-Shot Abductive ReasoningXiaoming Shi, Siqiao Xue, Kangrui Wang, Fan Zhou 等NeurIPS 2023 · 被引用 95 次
- Improving Time Sensitivity for Question Answering over Temporal Knowledge GraphsChao Shang, Guangtao Wang, Peng Qi, Jing HuangACL 2022 · 被引用 55 次
- Faking Fake News for Real Fake News Detection: Propaganda-Loaded Training Data GenerationKung-Hsiang Huang, Kathleen R. McKeown, Preslav Nakov, Yejin Choi 等ACL 2023 · 被引用 35 次
- ECONET: Effective Continual Pretraining of Language Models for Event Temporal ReasoningRujun Han, Xiang Ren, Nanyun PengEMNLP 2021 · 被引用 31 次
它引用的顶会 Paper2
相关 Paper
- Living in the Moment: Can Large Language Models Grasp Co-Temporal Reasoning?Zhaochen Su, Juntao Li, Jun Zhang, Tong Zhu 等ACL 2024
- Will LLMs Replace the Encoder-Only Models in Temporal Relation Classification?Gabriel Roccabruna, Massimo Rizzoli, Giuseppe RiccardiEMNLP 2024 · 被引用 3 次
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha 等EMNLP 2023 · 被引用 17 次
- ESTER: A Machine Reading Comprehension Dataset for Reasoning about Event Semantic RelationsRujun Han, I-Hung Hsu, Jiao Sun, Julia Baylon 等EMNLP 2021 · 被引用 30 次
- Analyzing Temporal Complex Events with Large Language Models? A Benchmark towards Temporal, Long Context UnderstandingZhihan Zhang, Yixin Cao, Chenchen Ye, Yunshan Ma 等ACL 2024
