Baleen: Robust Multi-Hop Reasoning at Scale via Condensed Retrieval
Omar Khattab, Christopher Potts, Matei A. Zaharia
摘要
Multi-hop reasoning (i.e., reasoning across two or more documents) is a key ingredient for NLP models that leverage large corpora to exhibit broad knowledge. To retrieve evidence passages, multi-hop models must contend with a fast-growing search space across the hops, represent complex queries that combine multiple information needs, and resolve ambiguity about the best order in which to hop between training passages. We tackle these problems via Baleen, a system that improves the accuracy of multi-hop retrieval while learning robustly from weak training signals in the many-hop setting. To tame the search space, we propose condensed retrieval, a pipeline that summarizes the retrieved passages after each hop into a single compact context. To model complex queries, we introduce a focused late interaction retriever that allows different parts of the same query representation to match disparate relevant passages. Lastly, to infer the hopping dependencies among unordered training passages, we devise latent hop ordering, a weak-supervision strategy in which the trained retriever itself selects the sequence of hops. We evaluate Baleen on retrieval for two-hop question answering and many-hop claim verification, establishing state-of-the-art performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalACL 2023 · 被引用 187 次
- DSPy: Compiling Declarative Language Model Calls into State-of-the-Art PipelinesOmar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang 等ICLR 2024 · 被引用 170 次
- PoET: A generative model of protein families as sequences-of-sequencesTimothy F. Truong Jr., Tristan BeplerNeurIPS 2023 · 被引用 96 次
- SuRe: Summarizing Retrievals using Answer Candidates for Open-domain QA of LLMsJaehyung Kim, Jaehyun Nam, Sangwoo Mo, Jongjin Park 等ICLR 2024 · 被引用 89 次
- Iteratively Prompt Pre-trained Language Models for Chain of ThoughtBoshi Wang, Xiang Deng, Huan SunEMNLP 2022 · 被引用 62 次
它引用的顶会 Paper7
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher 等ICLR 2020 · 被引用 322 次
- Answering Complex Open-Domain Questions with Multi-Hop Dense RetrievalWenhan Xiong, Xiang Lorraine Li, Srini Iyer, Jingfei Du 等ICLR 2021 · 被引用 232 次
相关 Paper
- Natural Logic-guided Autoregressive Multi-hop Document Retrieval for Fact VerificationRami Aly, Andreas VlachosEMNLP 2022 · 被引用 9 次
- HopRetriever: Retrieve Hops over Wikipedia to Answer Complex QuestionsShaobo Li, Xiaoguang Li, Lifeng Shang, Xin Jiang 等AAAI 2021 · 被引用 36 次
- BeamAggR: Beam Aggregation Reasoning over Multi-source Knowledge for Multi-hop Question AnsweringZheng Chu, Jingchang Chen, Qianglong Chen, Haotian Wang 等ACL 2024 · 被引用 8 次
- Coordinating Search-Informed Reasoning and Reasoning-Guided Search in Claim VerificationQisheng Hu, Quanyu Long, Wenya WangACL 2026 · 被引用 4 次
- Triple-Fact Retriever: An explainable reasoning retrieval model for multi-hop QA problemChengmin Wu, Enrui Hu, Ke Zhan, Lan Luo 等ICDE 2022 · 被引用 5 次
