What do Models Learn from Question Answering Datasets?
Priyanka Sen, Amir Saffari
Abstract
While models have reached superhuman performance on popular question answering (QA) datasets such as SQuAD, they have yet to outperform humans on the task of question answering itself. In this paper, we investigate if models are learning reading comprehension from QA datasets by evaluating BERT-based models across five datasets. We evaluate models on their generalizability to out-of-domain examples, responses to missing or incorrect data, and ability to handle question variations. We find that no single dataset is robust to all of our experiments and identify shortcomings in both datasets and evaluation methods. Following our analysis, we make recommendations for building future QA datasets that better evaluate the task of question answering through reading comprehension. We also release code to convert QA datasets to a shared format for easier experimentation at https: //github.com/amazon-research/ qa-dataset-converter.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 337 citations
- The Effect of Natural Distribution Shift on Question Answering ModelsJohn Miller, Karl Krauth, Benjamin Recht, Ludwig SchmidtICML 2020 · 158 citations
- CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about NegationAbhilasha Ravichander, Matt Gardner, Ana MarasovicEMNLP 2022 · 16 citations
- Weakly Supervised Neural Symbolic Learning for Cognitive TasksJidong Tian, Yitian Li, Wenqing Chen, Liqiang Xiao et al.AAAI 2022 · 14 citations
- IDK-MRC: Unanswerable Questions for Indonesian Machine Reading ComprehensionRifki Afina Putri, Alice OhEMNLP 2022 · 11 citations
Builds on1
Related papers
- Generative Language Models for Paragraph-Level Question GenerationAsahi Ushio, Fernando Alva-Manchego, José Camacho-ColladosEMNLP 2022 · 30 citations
- Span Selection Pre-training for Question AnsweringMichael R. Glass, Alfio Gliozzo, Rishav Chakravarti, Anthony Ferritto et al.ACL 2020 · 9 citations
- Assessing the Benchmarking Capacity of Machine Reading Comprehension DatasetsSaku Sugawara, Pontus Stenetorp, Kentaro Inui, Akiko AizawaAAAI 2020 · 92 citations
- A Robust Adversarial Training Approach to Machine Reading ComprehensionKai Liu, Xin Liu, An Yang, Jing Liu et al.AAAI 2020 · 54 citations
- Recurrent Chunking Mechanisms for Long-Text Machine Reading ComprehensionHongyu Gong, Yelong Shen, Dian Yu, Jianshu Chen et al.ACL 2020 · 39 citations
