The Distracting Effect: Understanding Irrelevant Passages in RAG
Chen Amiraz, Florin Cuconasu, Simone Filice, Zohar S. Karnin
Abstract
A well-known issue with Retrieval Augmented Generation (RAG) is that retrieved passages that are irrelevant to the query sometimes distract the answer-generating LLM, causing it to provide an incorrect response. In this paper, we shed light on this core issue and formulate the distracting effect of a passage w.r.t. a query (and an LLM). We provide a quantifiable measure of the distracting effect of a passage and demonstrate its robustness across LLMs. Our research introduces novel methods for identifying and using hard distracting passages to improve RAG systems. By fine-tuning LLMs with these carefully selected distracting passages, we achieve up to a 7.5% increase in answering accuracy compared to counterparts fine-tuned on conventional RAG datasets. Our contribution is two-fold: first, we move beyond the simple binary classification of irrelevant passages as either completely unrelated vs. distracting, and second, we develop and analyze multiple methods for finding hard distracting passages. To our knowledge, no other research has provided such a comprehensive framework for identifying and utilizing hard distracting passages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1374fb3-a57c-4221-a8fa-05e066a170dcCited by top-tier papers5
- Evaluating Legal Reasoning Traces with Legal Issue Tree RubricsJinu Lee, Kyoung-Woon On, Sophia Simeng Han, Arman Cohan et al.ACL 2026 · 2 citations
- Beyond RAG vs. Long-Context: Learning Distraction-Aware Retrieval for Efficient Knowledge GroundingSeong-Woong Shim, Myunsoo Kim, Jae Hyeon Cho, Byung-Jun LeeICLR 2026 · 1 citation
- Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAGXihang Wang, Zihan Wang, Chengkai Huang, Cao Liu et al.SIGIR 2026 · 1 citation
- Rethinking Soft Compression in Retrieval-Augmented Generation: A Query-Conditioned Selector PerspectiveYunhao Liu, Zian Jia, Xinyu Gao, Kanjun Xu et al.WWW 2026
- Efficient Prior-Guided Reasoning for Robust Retrieval-Augmented Generation under ConflictsXiaowei Yuan, Ziyang Huang, Zhao Yang, Yequan Wang et al.ACL 2026
Builds on9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das et al.ACL 2023 · 233 citations
- RA-DIT: Retrieval-Augmented Dual Instruction TuningXi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi et al.ICLR 2024 · 229 citations
Related papers
- Are Large Language Models Good at Utility Judgments?Hengran Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.SIGIR 2024 · 20 citations
- The Power of Noise: Redefining Retrieval for RAG SystemsFlorin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice et al.SIGIR 2024 · 212 citations
- Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAGBowen Jin, Jinsung Yoon, Jiawei Han, Sercan Ö. ArikICLR 2025
- GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal SynthesisYi Jiang, Sendong Zhao, Jianbo Li, Haochun Wang et al.ACL 2025
- RAGEval: Scenario Specific RAG Evaluation Dataset Generation FrameworkKunlun Zhu, Yifan Luo, Dingling Xu, Yukun Yan et al.ACL 2025 · 53 citations
