Answering Open-Domain Questions of Varying Reasoning Steps from Text
Peng Qi, Haejun Lee, Tg Sido, Christopher D. Manning
Abstract
We develop a unified system to answer directly from text open-domain questions that may require a varying number of retrieval steps. We employ a single multi-task transformer model to perform all the necessary subtasks-retrieving supporting facts, reranking them, and predicting the answer from all retrieved documents-in an iterative fashion. We avoid crucial assumptions of previous work that do not transfer well to real-world settings, including exploiting knowledge of the fixed number of retrieval steps required to answer each question or using structured metadata like knowledge bases or web links that have limited availability. Instead, we design a system that can answer open-domain questions on any text collection without prior knowledge of reasoning complexity. To emulate this setting, we construct a new benchmark, called B QA, by combining existing one-and twostep datasets with a new collection of 530 questions that require three Wikipedia pages to answer, unifying Wikipedia corpora versions in the process. We show that our model demonstrates competitive performance on both existing benchmarks and this new benchmark. We make the new benchmark available at https: //beerqa.github.io/. The Lord of the Rings à "150 million copies" Q. The Ingerophrynus Gollum is named after a character in a book that sold how many copies? Retriever Reader A. 150 million copies Answer exists in one of the reasoning paths No answer exist Repeat N times until the answer found is confident enough Reranker Query Generator Expand reasoning path with top-ranked paragraph … WIKIPEDIA search ① Q à "Ingerophrynus Gollum" ④ Q + Ingerophrynus Gollum à "Lord of the Rings" ② Q + retrieved paras à NOANSWER ⑤ Q + Ingerophrynus Gollum + The Lord of the Rings à "150 million copies"
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- DecAF: Joint Decoding of Answers and Logical Forms for Question Answering over Knowledge BasesDonghan Yu, Sheng Zhang, Patrick Ng, Henghui Zhu et al.ICLR 2023 · 18 citations
- Enhancing Multi-modal Multi-hop Question Answering via Structured Knowledge and Unified Retrieval-GenerationQian Yang, Qian Chen, Wen Wang, Baotian Hu et al.ACM MM 2023 · 14 citations
- Generative Multi-hop RetrievalHyunji Lee, Sohee Yang, Hanseok Oh, Minjoon SeoEMNLP 2022 · 12 citations
- Chain-of-Skills: A Configurable Model for Open-Domain Question AnsweringKaixin Ma, Hao Cheng, Yu Zhang, Xiaodong Liu et al.ACL 2023 · 12 citations
- Natural Logic-guided Autoregressive Multi-hop Document Retrieval for Fact VerificationRami Aly, Andreas VlachosEMNLP 2022 · 9 citations
Builds on8
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher et al.ICLR 2020 · 322 citations
- Answering Complex Open-Domain Questions with Multi-Hop Dense RetrievalWenhan Xiong, Xiang Lorraine Li, Srini Iyer, Jingfei Du et al.ICLR 2021 · 232 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Transformer-XH: Multi-Evidence Reasoning with eXtra Hop AttentionChen Zhao, Chenyan Xiong, Corby Rosset, Xia Song et al.ICLR 2020 · 120 citations
Related papers
- You Only Need One Model for Open-domain Question AnsweringHaejun Lee, Akhil Kedia, Jongwon Lee, Ashwin Paranjape et al.EMNLP 2022
- Answering Any-hop Open-domain Questions with Iterative Document RerankingYuyu Zhang, Ping Nie, Arun Ramamurthy, Le SongSIGIR 2021 · 18 citations
- Triple-Fact Retriever: An explainable reasoning retrieval model for multi-hop QA problemChengmin Wu, Enrui Hu, Ke Zhan, Lan Luo et al.ICDE 2022 · 5 citations
- Making Retrieval-Augmented Language Models Robust to Irrelevant ContextOri Yoran, Tomer Wolfson, Ori Ram, Jonathan BerantICLR 2024 · 361 citations
- CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmarkJian Wu, Linyi Yang, Zhen Wang, Manabu Okumura et al.ICLR 2025
