Challenges in Information-Seeking QA: Unanswerable Questions and Paragraph Retrieval
Akari Asai, Eunsol Choi
摘要
Recent pretrained language models "solved" many reading comprehension benchmarks, where questions are written with the access to the evidence document. However, datasets containing information-seeking queries where evidence documents are provided after the queries are written independently remain challenging. We analyze why answering information-seeking queries is more challenging and where their prevalent unanswerabilities arise, on Natural Questions and TyDi QA. Our controlled experiments suggest two headrooms -paragraph selection and answerability prediction, i.e. whether the paired evidence document contains the answer to the query or not. When provided with a gold paragraph and knowing when to abstain from answering, existing models easily outperform a human annotator. However, predicting answerability itself remains challenging. We manually annotate 800 unanswerable examples across six languages on what makes them challenging to answer. With this new data, we conduct percategory answerability prediction, revealing issues in the current dataset collection as well as task formulation. Together, our study points to avenues for future research in informationseeking question answering, both for dataset creation and model development. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- One Question Answering Model for Many Languages with Cross-lingual Dense Passage RetrievalAkari Asai, Xinyan Yu, Jungo Kasai, Hanna HajishirziNeurIPS 2021 · 被引用 86 次
- Improving Transformers with Probabilistic Attention KeysTam Minh Nguyen, Tan Minh Nguyen, Dung D. Le, Duy Khuong Nguyen 等ICML 2022 · 被引用 38 次
- Improving Transformer with an Admixture of Attention HeadsTan Nguyen, Tam Nguyen, Hai Do, Khai Nguyen 等NeurIPS 2022 · 被引用 38 次
- Towards Understanding Factual Knowledge of Large Language ModelsXuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo 等ICLR 2024 · 被引用 21 次
- DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question AnsweringElla Neeman, Roee Aharoni, Or Honovich, Leshem Choshen 等ACL 2023 · 被引用 20 次
它引用的顶会 Paper10
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Retrospective Reader for Machine Reading ComprehensionZhuosheng Zhang, Junjie Yang, Hai ZhaoAAAI 2021 · 被引用 237 次
- MultiModalQA: complex question answering over text, tables and imagesAlon Talmor, Ori Yoran, Amnon Catav, Dan Lahav 等ICLR 2021 · 被引用 229 次
- SG-Net: Syntax-Guided Machine Reading ComprehensionZhuosheng Zhang, Yuwei Wu, Junru Zhou, Sufeng Duan 等AAAI 2020 · 被引用 192 次
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 被引用 162 次
相关 Paper
- Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun 等EMNLP 2023 · 被引用 37 次
- Do I have the Knowledge to Answer? Investigating Answerability of Knowledge Base QuestionsMayur Patidar, Prayushi Faldu, Avinash Kumar Singh, Lovekesh Vig 等ACL 2023 · 被引用 2 次
- (QA)²: Question Answering with Questionable AssumptionsNajoung Kim, Phu Mon Htut, Samuel R. Bowman, Jackson PettyACL 2023 · 被引用 2 次
- Inquisitive Question Generation for High Level Text ComprehensionWei-Jen Ko, Te-Yuan Chen, Yiyan Huang, Greg Durrett 等EMNLP 2020 · 被引用 32 次
- IIRC: A Dataset of Incomplete Information Reading Comprehension QuestionsJames Ferguson, Matt Gardner, Hannaneh Hajishirzi, Tushar Khot 等EMNLP 2020 · 被引用 42 次
