What do Models Learn from Question Answering Datasets?
Priyanka Sen, Amir Saffari
摘要
While models have reached superhuman performance on popular question answering (QA) datasets such as SQuAD, they have yet to outperform humans on the task of question answering itself. In this paper, we investigate if models are learning reading comprehension from QA datasets by evaluating BERT-based models across five datasets. We evaluate models on their generalizability to out-of-domain examples, responses to missing or incorrect data, and ability to handle question variations. We find that no single dataset is robust to all of our experiments and identify shortcomings in both datasets and evaluation methods. Following our analysis, we make recommendations for building future QA datasets that better evaluate the task of question answering through reading comprehension. We also release code to convert QA datasets to a shared format for easier experimentation at https: //github.com/amazon-research/ qa-dataset-converter.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 被引用 337 次
- The Effect of Natural Distribution Shift on Question Answering ModelsJohn Miller, Karl Krauth, Benjamin Recht, Ludwig SchmidtICML 2020 · 被引用 158 次
- CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about NegationAbhilasha Ravichander, Matt Gardner, Ana MarasovicEMNLP 2022 · 被引用 16 次
- Weakly Supervised Neural Symbolic Learning for Cognitive TasksJidong Tian, Yitian Li, Wenqing Chen, Liqiang Xiao 等AAAI 2022 · 被引用 14 次
- IDK-MRC: Unanswerable Questions for Indonesian Machine Reading ComprehensionRifki Afina Putri, Alice OhEMNLP 2022 · 被引用 11 次
它引用的顶会 Paper1
相关 Paper
- Generative Language Models for Paragraph-Level Question GenerationAsahi Ushio, Fernando Alva-Manchego, José Camacho-ColladosEMNLP 2022 · 被引用 30 次
- Span Selection Pre-training for Question AnsweringMichael R. Glass, Alfio Gliozzo, Rishav Chakravarti, Anthony Ferritto 等ACL 2020 · 被引用 9 次
- Assessing the Benchmarking Capacity of Machine Reading Comprehension DatasetsSaku Sugawara, Pontus Stenetorp, Kentaro Inui, Akiko AizawaAAAI 2020 · 被引用 92 次
- A Robust Adversarial Training Approach to Machine Reading ComprehensionKai Liu, Xin Liu, An Yang, Jing Liu 等AAAI 2020 · 被引用 54 次
- Recurrent Chunking Mechanisms for Long-Text Machine Reading ComprehensionHongyu Gong, Yelong Shen, Dian Yu, Jianshu Chen 等ACL 2020 · 被引用 39 次
