Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
Shiping Yang, Jie Wu, Wenbiao Ding, Ning Wu, Shining Liang, Ming Gong, Hongzhi Li, Hengyuan Zhang, Angel X. Chang, Dongmei Zhang
Abstract
Robustness has become a critical attribute for the deployment of RAG systems in real-world applications. Existing research focuses on robustness to explicit noise (e.g., document semantics) but overlooks implicit noise (spurious features). Moreover, previous studies on spurious features in LLMs are limited to specific types (e.g., formats) and narrow scenarios (e.g., ICL). In this work, we identify and study spurious features in the RAG paradigm, a robustness issue caused by the sensitivity of LLMs to semantic-agnostic features. We then propose a novel framework, SURE, to empirically quantify the robustness of RALMs against spurious features. Beyond providing a comprehensive taxonomy and metrics for evaluation, the framework's data synthesis pipeline facilitates training-based strategies to improve robustness. Further analysis suggests that spurious features are a widespread and challenging problem in the field of RAG. Our code is available at https://github.com/maybenotime/ RAG-SpuriousFeatures .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 682 citations
- The Power of Noise: Redefining Retrieval for RAG SystemsFlorin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice et al.SIGIR 2024 · 212 citations
- Retrieval meets Long Context Large Language ModelsPeng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee et al.ICLR 2024 · 131 citations
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi et al.EMNLP 2024 · 119 citations
Related papers
- SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language ModelXun Liang, Simin Niu, Zhiyu Li, Sensen Zhang et al.ACL 2025
- PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented GenerationZhehao Tan, Yihan Jiao, Dan Yang, Junwei Liu et al.AAAI 2026
- RAGEval: Scenario Specific RAG Evaluation Dataset Generation FrameworkKunlun Zhu, Yifan Luo, Dingling Xu, Yukun Yan et al.ACL 2025 · 53 citations
- RAGGED: Towards Informed Design of Scalable and Stable RAG SystemsJennifer Hsia, Afreen Shaikh, Zora Zhiruo Wang, Graham NeubigICML 2025
- SURE or Not? Investigating Semantic Understanding in Dense Retrieval ModelsLingdi Kong, Xuanang Chen, Ben He, Le SunACL 2026
