Spiral of Silence: How is Large Language Model Killing Information Retrieval? - A Case Study on Open Domain Question Answering
Xiaoyang Chen, Ben He, Hongyu Lin, Xianpei Han, Tianshu Wang, Boxi Cao, Le Sun, Yingfei Sun
Abstract
The practice of Retrieval-Augmented Generation (RAG), which integrates Large Language Models (LLMs) with retrieval systems, has become increasingly prevalent. However, the repercussions of LLM-derived content infiltrating the web and influencing the retrievalgeneration feedback loop are largely uncharted territories. In this study, we construct and iteratively run a simulation pipeline to deeply investigate the short-term and long-term effects of LLM text on RAG systems. Taking the trending Open Domain Question Answering (ODQA) task as a point of entry, our findings reveal a potential digital "Spiral of Silence" effect, with LLM-generated text consistently outperforming human-authored content in search rankings, thereby diminishing the presence and impact of human contributions online. This trend risks creating an imbalanced information ecosystem, where the unchecked proliferation of erroneous LLM-generated content may result in the marginalization of accurate information. We urge the academic community to take heed of this potential issue, ensuring a diverse and authentic digital information landscape. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- TFRank: Think-Free Reasoning Enables Practical Pointwise LLM RankingYongqi Fan, Xiaoyang Chen, Dezhi Ye, Jie Liu et al.AAAI 2026 · 9 citations
- LLM-Generated Fake News Induces Truth Decay in News Ecosystem: A Case Study on Neural News RecommendationBeizhe Hu, Qiang Sheng, Juan Cao, Yang Li et al.SIGIR 2025 · 7 citations
- Enhancing LLM Text Detection with Retrieved Contexts and Logits Distribution ConsistencyZhaoheng Huang, Yutao Zhu, Ji-Rong Wen, Zhicheng DouEMNLP 2025
- LLM-Generated Text May Harm Your Retrieval! A Robust Detection Strategy for Retrieval-Augmented GenerationZhaoheng Huang, Yutao Zhu, Ji-Rong Wen, Zhicheng DouACL 2026
Builds on14
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 531 citations
Related papers
- Information Retrieval in the Age of Generative AI: The RGB ModelMichele Garetto, Alessandro Cornacchia, Franco Galante, Emilio Leonardi et al.SIGIR 2025 · 2 citations
- The Power of Noise: Redefining Retrieval for RAG SystemsFlorin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice et al.SIGIR 2024 · 212 citations
- Exploring the Escalation of Source Bias in User, Data, and Recommender System Feedback LoopYuqi Zhou, Sunhao Dai, Liang Pang, Gang Wang et al.SIGIR 2025 · 2 citations
- Rethinking the Hidden Risk of Reranking: Achieving Risk-aware Reranking with Information Gain for RAG with LLMsZhizhao Liu, Zhihua Wen, Zhiliang Tian, Zhen Huang et al.WWW 2026
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
