Answering Open-Domain Multi-Answer Questions via a Recall-then-Verify Framework
Zhihong Shao, Minlie Huang
Abstract
Open-domain questions are likely to be openended and ambiguous, leading to multiple valid answers. Existing approaches typically adopt the rerank-then-read framework, where a reader reads top-ranking evidence to predict answers. According to our empirical analysis, this framework faces three problems: first, to leverage a large reader under a memory constraint, the reranker should select only a few relevant passages to cover diverse answers, while balancing relevance and diversity is non-trivial; second, the small reading budget prevents the reader from accessing valuable retrieved evidence filtered out by the reranker; third, when using a generative reader to predict answers all at once based on all selected evidence, whether a valid answer will be predicted also pathologically depends on the evidence of some other valid answer(s). To address these issues, we propose to answer open-domain multi-answer questions with a recall-then-verify framework, which separates the reasoning process of each answer so that we can make better use of retrieved evidence while also leveraging large models under the same memory constraint. Our framework achieves state-of-the-art results on two multi-answer datasets, and predicts significantly more gold answers than a rerank-thenread system that uses an oracle reranker.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 550cb0da-496b-497b-aa9b-3af1ae128779Cited by top-tier papers4
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen et al.ICLR 2024 · 699 citations
- BlendFilter: Advancing Retrieval-Augmented Large Language Models via Query Generation Blending and Knowledge FilteringHaoyu Wang, Ruirui Li, Haoming Jiang, Jinjin Tian et al.EMNLP 2024 · 9 citations
- Answering Ambiguous Questions via Iterative PromptingWeiwei Sun, Hengyi Cai, Hongshen Chen, Pengjie Ren et al.ACL 2023 · 3 citations
- Adaptive Question Answering: Enhancing Language Model Proficiency for Addressing Knowledge Conflicts with Source CitationsSagi Shaier, Ari Kobren, Philip V. OgrenEMNLP 2024
Builds on9
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Distilling Knowledge from Reader to Retriever for Question AnsweringGautier Izacard, Edouard GraveICLR 2021 · 317 citations
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 162 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- Answering Any-hop Open-domain Questions with Iterative Document RerankingYuyu Zhang, Ping Nie, Arun Ramamurthy, Le SongSIGIR 2021 · 18 citations
- Joint Passage Ranking for Diverse Multi-Answer RetrievalSewon Min, Kenton Lee, Ming-Wei Chang, Kristina Toutanova et al.EMNLP 2021 · 1 citation
- Harnessing Multi-Role Capabilities of Large Language Models for Open-Domain Question AnsweringHongda Sun, Yuxuan Liu, Chengwei Wu, Haiyu Yan et al.WWW 2024 · 16 citations
- You Only Need One Model for Open-domain Question AnsweringHaejun Lee, Akhil Kedia, Jongwon Lee, Ashwin Paranjape et al.EMNLP 2022
- RINK: Reader-Inherited Evidence Reranker for Table-and-Text Open Domain Question AnsweringEunhwan Park, Sung-Min Lee, Daeryong Seo, Seonhoon Kim et al.AAAI 2023 · 4 citations
