Neural Retrievers are Biased Towards LLM-Generated Content
Sunhao Dai, Yuqi Zhou, Liang Pang, Weihao Liu, Xiaolin Hu, Yong Liu, Xiao Zhang, Gang Wang, Jun Xu
Abstract
Recently, the emergence of large language models (LLMs) has revolutionized the paradigm of information retrieval (IR) applications, especially in web search, by generating vast amounts of human-like texts on the Internet. As a result, IR systems in the LLM era are facing a new challenge: the indexed documents are now not only written by human beings but also automatically generated by the LLMs. How these LLM-generated documents influence the IR systems is a pressing and still unexplored question. In this work, we conduct a quantitative evaluation of IR models in scenarios where both human-written and LLM-generated texts are involved. Surprisingly, our findings indicate that neural retrieval models tend to rank LLM-generated documents higher. We refer to this category of biases in neural retrievers towards the LLM-generated content as the source bias. Moreover, we discover that this bias is not confined to the first-stage neural retrievers, but extends to the second-stage neural re-rankers. Then, in-depth analyses from the perspective of text compression indicate that LLM-generated texts exhibit more focused semantics with less noise, making it easier for neural retrieval models to semantic match. To mitigate the source bias, we also propose a plug-and-play debiased constraint for the optimization objective, and experimental results show its effectiveness. Finally, we discuss the potential severe concerns stemming from the observed source bias and hope our findings can serve as a critical wake-up call to the IR community and beyond. To facilitate future explorations of IR in the LLM era, the constructed two new benchmarks are available at https://github.com/KID-22/Source-Bias.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3db5bca0-709f-49ad-9467-1be79f971733Cited by top-tier papers14
- Invisible Relevance Bias: Text-Image Retrieval Models Prefer AI-Generated ImagesShicheng Xu, Danyang Hou, Liang Pang, Jingcheng Deng et al.SIGIR 2024 · 18 citations
- Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-in-the-Loop LLMJiachen Li, Xiwen Li, Justin Steinberg, Akshat Choube et al.UbiComp 2025 · 9 citations
- Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding DataShiping Yang, Jie Wu, Wenbiao Ding, Ning Wu et al.ACL 2026 · 9 citations
- LLM-Generated Fake News Induces Truth Decay in News Ecosystem: A Case Study on Neural News RecommendationBeizhe Hu, Qiang Sheng, Juan Cao, Yang Li et al.SIGIR 2025 · 7 citations
- AIGCs Confuse AI Too: Investigating and Explaining Synthetic Image-induced Hallucinations in Large Vision-Language ModelsYifei Gao, Jiaqi Wang, Zhiyu Lin, Jitao SangACM MM 2024 · 7 citations
Builds on8
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin et al.SIGIR 2021 · 297 citations
- Self-Consuming Generative Models Go MADSina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun et al.ICLR 2024 · 279 citations
- Language Modeling Is CompressionGrégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt et al.ICLR 2024 · 243 citations
Related papers
- Mitigating Source Bias with LLM AlignmentSunhao Dai, Yuqi Zhou, Liang Pang, Zhuoyang Li et al.SIGIR 2025 · 2 citations
- Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity DocumentsHaoyu Wang, Sunhao Dai, Haiyuan Zhao, Liang Pang et al.ICLR 2025
- Exploring the Escalation of Source Bias in User, Data, and Recommender System Feedback LoopYuqi Zhou, Sunhao Dai, Liang Pang, Gang Wang et al.SIGIR 2025 · 2 citations
- AIR-Bench: Automated Heterogeneous Information Retrieval BenchmarkJianlyu Chen, Nan Wang, Chaofan Li, Bo Wang et al.ACL 2025
- Attention in Large Language Models Yields Efficient Zero-Shot Re-RankersShijie Chen, Bernal Jimenez Gutierrez, Yu SuICLR 2025
