English Machine Reading Comprehension Datasets: A Survey
Daria Dzendzik, Jennifer Foster, Carl Vogel
Abstract
This paper surveys 60 English Machine Reading Comprehension datasets, with a view to providing a convenient resource for other researchers interested in this problem. We categorize the datasets according to their question and answer form and compare them across various dimensions including size, vocabulary, data source, method of creation, human performance level, and first question word. Our analysis reveals that Wikipedia is by far the most common data source and that there is a relative lack of why, when, and where questions across datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb767a73-7ab1-42da-ad49-564c037d8d98Cited by top-tier papers5
- Improving Unsupervised Question Answering via Summarization-Informed Question GenerationChenyang Lyu, Lifeng Shang, Yvette Graham, Jennifer Foster et al.EMNLP 2021 · 33 citations
- VASR: Visual Analogies of Situation RecognitionYonatan Bitton, Ron Yosef, Eliyahu Strugo, Dafna Shahaf et al.AAAI 2023 · 27 citations
- Distractor Generation in Multiple-Choice Tasks: A Survey of Methods, Datasets, and EvaluationElaf Alhazmi, Quan Sheng, Wei Emma Zhang, Munazza Zaib et al.EMNLP 2024 · 15 citations
- Clickbait Spoiling via Question Answering and Passage RetrievalMatthias Hagen, Maik Fröbe, Artur Jurk, Martin PotthastACL 2022
- A Comprehensive Literary Chinese Reading Comprehension Dataset with an Evidence Curation Based SolutionDongning Rao, Rongchu Zhou, Peng Chen, Zhihua JiangEMNLP 2025
Builds on9
- QASC: A Dataset for Question Answering via Sentence CompositionTushar Khot, Peter Clark, Michal Guerquin, Peter Jansen et al.AAAI 2020 · 387 citations
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 325 citations
- Getting Closer to AI Complete Question Answering: A Set of Prerequisite Real TasksAnna Rogers, Olga Kovaleva, Matthew Downey, Anna RumshiskyAAAI 2020 · 141 citations
- MLQA: Evaluating Cross-lingual Extractive Question AnsweringPatrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel et al.ACL 2020 · 52 citations
- IIRC: A Dataset of Incomplete Information Reading Comprehension QuestionsJames Ferguson, Matt Gardner, Hannaneh Hajishirzi, Tushar Khot et al.EMNLP 2020 · 42 citations
Related papers
- WebSRC: A Dataset for Web-Based Structural Reading ComprehensionXingyu Chen, Zihan Zhao, Lu Chen, Jiabao Ji et al.EMNLP 2021 · 42 citations
- What do Models Learn from Question Answering Datasets?Priyanka Sen, Amir SaffariEMNLP 2020 · 40 citations
- VisualMRC: Machine Reading Comprehension on Document ImagesRyota Tanaka, Kyosuke Nishida, Sen YoshidaAAAI 2021 · 201 citations
- Assessing the Benchmarking Capacity of Machine Reading Comprehension DatasetsSaku Sugawara, Pontus Stenetorp, Kentaro Inui, Akiko AizawaAAAI 2020 · 92 citations
- Inquisitive Question Generation for High Level Text ComprehensionWei-Jen Ko, Te-Yuan Chen, Yiyan Huang, Greg Durrett et al.EMNLP 2020 · 32 citations
