Challenges in Information-Seeking QA: Unanswerable Questions and Paragraph Retrieval
Akari Asai, Eunsol Choi
Abstract
Recent pretrained language models "solved" many reading comprehension benchmarks, where questions are written with the access to the evidence document. However, datasets containing information-seeking queries where evidence documents are provided after the queries are written independently remain challenging. We analyze why answering information-seeking queries is more challenging and where their prevalent unanswerabilities arise, on Natural Questions and TyDi QA. Our controlled experiments suggest two headrooms -paragraph selection and answerability prediction, i.e. whether the paired evidence document contains the answer to the query or not. When provided with a gold paragraph and knowing when to abstain from answering, existing models easily outperform a human annotator. However, predicting answerability itself remains challenging. We manually annotate 800 unanswerable examples across six languages on what makes them challenging to answer. With this new data, we conduct percategory answerability prediction, revealing issues in the current dataset collection as well as task formulation. Together, our study points to avenues for future research in informationseeking question answering, both for dataset creation and model development. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- One Question Answering Model for Many Languages with Cross-lingual Dense Passage RetrievalAkari Asai, Xinyan Yu, Jungo Kasai, Hanna HajishirziNeurIPS 2021 · 86 citations
- Improving Transformers with Probabilistic Attention KeysTam Minh Nguyen, Tan Minh Nguyen, Dung D. Le, Duy Khuong Nguyen et al.ICML 2022 · 38 citations
- Improving Transformer with an Admixture of Attention HeadsTan Nguyen, Tam Nguyen, Hai Do, Khai Nguyen et al.NeurIPS 2022 · 38 citations
- Towards Understanding Factual Knowledge of Large Language ModelsXuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo et al.ICLR 2024 · 21 citations
- DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question AnsweringElla Neeman, Roee Aharoni, Or Honovich, Leshem Choshen et al.ACL 2023 · 20 citations
Builds on10
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrospective Reader for Machine Reading ComprehensionZhuosheng Zhang, Junjie Yang, Hai ZhaoAAAI 2021 · 237 citations
- MultiModalQA: complex question answering over text, tables and imagesAlon Talmor, Ori Yoran, Amnon Catav, Dan Lahav et al.ICLR 2021 · 229 citations
- SG-Net: Syntax-Guided Machine Reading ComprehensionZhuosheng Zhang, Yuwei Wu, Junru Zhou, Sufeng Duan et al.AAAI 2020 · 192 citations
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 162 citations
Related papers
- Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun et al.EMNLP 2023 · 37 citations
- Do I have the Knowledge to Answer? Investigating Answerability of Knowledge Base QuestionsMayur Patidar, Prayushi Faldu, Avinash Kumar Singh, Lovekesh Vig et al.ACL 2023 · 2 citations
- (QA)²: Question Answering with Questionable AssumptionsNajoung Kim, Phu Mon Htut, Samuel R. Bowman, Jackson PettyACL 2023 · 2 citations
- Inquisitive Question Generation for High Level Text ComprehensionWei-Jen Ko, Te-Yuan Chen, Yiyan Huang, Greg Durrett et al.EMNLP 2020 · 32 citations
- IIRC: A Dataset of Incomplete Information Reading Comprehension QuestionsJames Ferguson, Matt Gardner, Hannaneh Hajishirzi, Tushar Khot et al.EMNLP 2020 · 42 citations
