Mitigating False-Negative Contexts in Multi-document Question Answering with Retrieval Marginalization
Ansong Ni, Matt Gardner, Pradeep Dasigi
Abstract
Question Answering (QA) tasks requiring information from multiple documents often rely on a retrieval model to identify relevant information for reasoning. The retrieval model is typically trained to maximize the likelihood of the labeled supporting evidence. However, when retrieving from large text corpora such as Wikipedia, the correct answer can often be obtained from multiple evidence candidates. Moreover, not all such candidates are labeled as positive during annotation, rendering the training signal weak and noisy. This problem is exacerbated when the questions are unanswerable or when the answers are Boolean, since the model cannot rely on lexical overlap to make a connection between the answer and supporting evidence. We develop a new parameterization of set-valued retrieval that handles unanswerable queries, and we show that marginalizing over this set during training allows a model to mitigate false negatives in supporting evidence annotations. We test our method on two multi-document QA datasets, IIRC and HotpotQA. On IIRC, we show that joint modeling with marginalization improves model performance by 5.5 F1 points and achieves a new state-of-the-art performance of 50.5 F1. We also show that retrieval marginalization results in 4.1 QA F1 improvement over a non-marginalized baseline on HotpotQA in the fullwiki setting. 1 * Majority of the work done as an intern at AI2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e83cae73-c62c-42f0-acd2-1a4efb306be6Cited by top-tier papers5
- Turning Tables: Generating Examples from Semi-structured Tables for Endowing Language Models with Reasoning SkillsOri Yoran, Alon Talmor, Jonathan BerantACL 2022 · 57 citations
- DSAS: A Universal Plug-and-Play Framework for Attention Optimization in Multi-Document Question AnsweringJiakai Li, Rongzheng Wang, Yizhuo Ma, Shuang Liang et al.NeurIPS 2025 · 8 citations
- Teaching Broad Reasoning Skills for Multi-Step QA by Generating Hard ContextsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalEMNLP 2022 · 6 citations
- Mitigating the Impact of False Negative in Dense Retrieval with Contrastive Confidence RegularizationShiqi Wang, Yeqin Zhang, Cam-Tu NguyenAAAI 2024 · 6 citations
- ARHN: Answer-Centric Relabeling of Hard Negatives with Open-Source LLMs for Dense RetrievalHyewon Choi, Jooyoung Choi, Hansol Jang, Hyun Kim et al.SIGIR 2026
Builds on6
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher et al.ICLR 2020 · 322 citations
- Hierarchical Graph Network for Multi-hop Question AnsweringYuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai et al.EMNLP 2020 · 157 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Neural Module Networks for Reasoning over TextNitish Gupta, Kevin Lin, Dan Roth, Sameer Singh et al.ICLR 2020 · 134 citations
- Turning Tables: Generating Examples from Semi-structured Tables for Endowing Language Models with Reasoning SkillsOri Yoran, Alon Talmor, Jonathan BerantACL 2022 · 57 citations
Related papers
- HopRetriever: Retrieve Hops over Wikipedia to Answer Complex QuestionsShaobo Li, Xiaoguang Li, Lifeng Shang, Xin Jiang et al.AAAI 2021 · 36 citations
- End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question AnsweringDevendra Singh Sachan, Siva Reddy, William L. Hamilton, Chris Dyer et al.NeurIPS 2021 · 197 citations
- Triple-Fact Retriever: An explainable reasoning retrieval model for multi-hop QA problemChengmin Wu, Enrui Hu, Ke Zhan, Lan Luo et al.ICDE 2022 · 5 citations
- Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented GenerationHengran Zhang, Minghao Tang, Keping Bi, Jiafeng Guo et al.EMNLP 2025 · 1 citation
- IIRC: A Dataset of Incomplete Information Reading Comprehension QuestionsJames Ferguson, Matt Gardner, Hannaneh Hajishirzi, Tushar Khot et al.EMNLP 2020 · 42 citations
