Getting Closer to AI Complete Question Answering: A Set of Prerequisite Real Tasks
Anna Rogers, Olga Kovaleva, Matthew Downey, Anna Rumshisky
Abstract
The recent explosion in question answering research produced a wealth of both factoid reading comprehension (RC) and commonsense reasoning datasets. Combining them presents a different kind of task: deciding not simply whether information is present in the text, but also whether a confident guess could be made for the missing information. We present QuAIL, the first RC dataset to combine text-based, world knowledge and unanswerable questions, and to provide question type annotation that would enable diagnostics of the reasoning strategies by a given QA system. QuAIL contains 15K multi-choice questions for 800 texts in 4 domains. Crucially, it offers both general and text-specific questions, unlikely to be found in pretraining data. We show that QuAIL poses substantial challenges to the current state-of-the-art systems, with a 30% drop in accuracy compared to the most similar existing dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3e49c25-72aa-4c8d-b4b1-0ce97fae6e62Cited by top-tier papers47
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker et al.ICML 2024 · 443 citations
- RA-DIT: Retrieval-Augmented Dual Instruction TuningXi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi et al.ICLR 2024 · 229 citations
- xRAG: Extreme Context Compression for Retrieval-augmented Generation with One TokenXin Cheng, Xun Wang, Xingxing Zhang, Tao Ge et al.NeurIPS 2024 · 156 citations
Related papers
- QASC: A Dataset for Question Answering via Sentence CompositionTushar Khot, Peter Clark, Michal Guerquin, Peter Jansen et al.AAAI 2020 · 387 citations
- Inferential Question AnsweringJamshid Mozafari, Hamed Zamani, Guido Zuccon, Adam JatowtWWW 2026
- WikiWhy: Answering and Explaining Cause-and-Effect QuestionsMatthew Ho, Aditya Sharma, Justin Chang, Michael Saxon et al.ICLR 2023 · 8 citations
- IIRC: A Dataset of Incomplete Information Reading Comprehension QuestionsJames Ferguson, Matt Gardner, Hannaneh Hajishirzi, Tushar Khot et al.EMNLP 2020 · 42 citations
- QUITE: Quantifying Uncertainty in Natural Language Text in Bayesian Reasoning ScenariosTimo Pierre Schrader, Lukas Lange, Simon Razniewski, Annemarie FriedrichEMNLP 2024
