Shifting from Ranking to Set Selection for Retrieval Augmented Generation
Dahyun Lee, Yongrae Jo, Haeju Park, Moontae Lee
Abstract
Retrieval in Retrieval-Augmented Generation (RAG) must ensure that retrieved passages are not only individually relevant but also collectively form a comprehensive set. Existing approaches primarily rerank top-k passages based on their individual relevance, often failing to meet the information needs of complex queries in multi-hop question answering. In this work, we propose a set-wise passage selection approach and introduce SETR, which explicitly identifies the information requirements of a query through Chain-of-Thought reasoning and selects an optimal set of passages that collectively satisfy those requirements. Experiments on multi-hop RAG benchmarks show that SETR outperforms both proprietary LLM-based rerankers and open-source baselines in terms of answer correctness and retrieval quality, providing an effective and efficient alternative to traditional rerankers in RAG systems. The code is available at https:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive DomainsYash Saxena, Ankur Padia, Mandar Chaudhary, Kalpa Gunaratna et al.ICML 2026 · 7 citations
- Beyond RAG vs. Long-Context: Learning Distraction-Aware Retrieval for Efficient Knowledge GroundingSeong-Woong Shim, Myunsoo Kim, Jae Hyeon Cho, Byung-Jun LeeICLR 2026 · 1 citation
- Search for Coverage: Learning Coverage-Aware Retrieval with Augmented Sub-Question AnswerabilityJia-Huei Ju, Eugene Yang, Trevor Adriaanse, Suzan Verberne et al.SIGIR 2026
- Why Retrieval-Augmented Generation Fails: A Graph PerspectiveKai Guo, Xinnan Dai, Zhibo Zhang, Nuohan Lin et al.KDD 2026
- FedMosaic: Federated Retrieval-Augmented Generation via Parametric AdaptersZhilin Liang, Yuxiang Wang, Zimu Zhou, Hainan Zhang et al.SIGIR 2026
Builds on10
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- RAPTOR: Recursive Abstractive Processing for Tree-Organized RetrievalParth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna et al.ICLR 2024 · 460 citations
Related papers
- Optimizing Question Semantic Space for Dynamic Retrieval-Augmented Multi-hop Question AnsweringLinhao Ye, Lang Yu, Zhikai Lei, Qin Chen et al.ACL 2025 · 4 citations
- Retriever Portfolios: A Principled Approach to Adaptive RAGMiltiadis Stouras, Vincent Cohen-Addad, Silvio Lattanzi, Ola SvenssonICML 2026
- Are Large Language Models Good at Utility Judgments?Hengran Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.SIGIR 2024 · 20 citations
- Iterative Multi-Granular RAG with Contextual Hierarchical GraphYanli Hu, Teng Liu, Zhuangyi Zhou, Weixin Zeng et al.AAAI 2026
- S2G-RAG: Structured Sufficiency and Gap Judging for Iterative Retrieval-Augmented QAMinghan Li, Junjie Zou, Xinxuan Lv, Chao Zhang et al.ACL 2026 · 1 citation
