Information Retrieval Induced Safety Degradation in AI Agents
Cheng Yu, Benedikt Stroebl, Diyi Yang, Orestis Papakyriakopoulos
Abstract
Despite the growing integration of retrieval-enabled AI agents into society, their safety and ethical behavior remain inadequately understood. In particular, the growing integration of LLMs and AI agents with external information sources and real-world environments raises critical questions about how they engage with and are influenced by these external data sources and interactive contexts. This study investigates how expanding retrieval access—from no external sources to Wikipedia-based retrieval and open web search—affects model reliability, bias propagation, and harmful content generation. Through extensive benchmarking of censored and uncensored LLMs and AI Agents, our findings reveal a consistent degradation in refusal rates, bias sensitivity, and harmfulness safeguards as models gain broader access to external sources, culminating in a phenomenon we term safety degradation . Notably, retrieval-enabled agents built on aligned LLMs often behave more unsafely than uncensored models without retrieval. This effect persists even under strong retrieval accuracy and prompt-based mitigation, suggesting that the mere presence of retrieved content reshapes model behavior in structurally unsafe ways. These findings underscore the need for robust mitigation strategies to ensure fairness and reliability in retrieval-enabled and increasingly autonomous AI systems. Content Warning : This paper contains examples of harmful language.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
Related papers
- ShieldRAG: Safeguarding Retrieval-Augmented Generation from Untrusted Knowledge BasesPeiru Yang, Haoran Zheng, Yi Luo, Xinyi Liu et al.AAAI 2026
- PurifAI: Detecting and Fixing Search-Induced Distortions in Web-Augmented LLMsGuoqing Wang, Zhao Zhang, Zeyu Sun, Xiaofei Xie et al.SIGIR 2026
- Jailbreak Open-Sourced Large Language Models via Enforced DecodingHangfan Zhang, Zhimeng Guo, Huaisheng Zhu, Bochuan Cao et al.ACL 2024
- Neural Retrievers are Biased Towards LLM-Generated ContentSunhao Dai, Yuqi Zhou, Liang Pang, Weihao Liu et al.KDD 2024 · 26 citations
- On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI AlignmentSarah Ball, Greg Gluch, Shafi Goldwasser, Frauke Kreuter et al.ICLR 2026 · 16 citations
