Neural Retrievers are Biased Towards LLM-Generated Content
Sunhao Dai, Yuqi Zhou, Liang Pang, Weihao Liu, Xiaolin Hu, Yong Liu, Xiao Zhang, Gang Wang, Jun Xu
摘要
Recently, the emergence of large language models (LLMs) has revolutionized the paradigm of information retrieval (IR) applications, especially in web search, by generating vast amounts of human-like texts on the Internet. As a result, IR systems in the LLM era are facing a new challenge: the indexed documents are now not only written by human beings but also automatically generated by the LLMs. How these LLM-generated documents influence the IR systems is a pressing and still unexplored question. In this work, we conduct a quantitative evaluation of IR models in scenarios where both human-written and LLM-generated texts are involved. Surprisingly, our findings indicate that neural retrieval models tend to rank LLM-generated documents higher. We refer to this category of biases in neural retrievers towards the LLM-generated content as the source bias. Moreover, we discover that this bias is not confined to the first-stage neural retrievers, but extends to the second-stage neural re-rankers. Then, in-depth analyses from the perspective of text compression indicate that LLM-generated texts exhibit more focused semantics with less noise, making it easier for neural retrieval models to semantic match. To mitigate the source bias, we also propose a plug-and-play debiased constraint for the optimization objective, and experimental results show its effectiveness. Finally, we discuss the potential severe concerns stemming from the observed source bias and hope our findings can serve as a critical wake-up call to the IR community and beyond. To facilitate future explorations of IR in the LLM era, the constructed two new benchmarks are available at https://github.com/KID-22/Source-Bias.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Invisible Relevance Bias: Text-Image Retrieval Models Prefer AI-Generated ImagesShicheng Xu, Danyang Hou, Liang Pang, Jingcheng Deng 等SIGIR 2024 · 被引用 18 次
- Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-in-the-Loop LLMJiachen Li, Xiwen Li, Justin Steinberg, Akshat Choube 等UbiComp 2025 · 被引用 9 次
- Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding DataShiping Yang, Jie Wu, Wenbiao Ding, Ning Wu 等ACL 2026 · 被引用 9 次
- LLM-Generated Fake News Induces Truth Decay in News Ecosystem: A Case Study on Neural News RecommendationBeizhe Hu, Qiang Sheng, Juan Cao, Yang Li 等SIGIR 2025 · 被引用 7 次
- AIGCs Confuse AI Too: Investigating and Explaining Synthetic Image-induced Hallucinations in Large Vision-Language ModelsYifei Gao, Jiaqi Wang, Zhiyu Lin, Jitao SangACM MM 2024 · 被引用 7 次
它引用的顶会 Paper8
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin 等SIGIR 2021 · 被引用 297 次
- Self-Consuming Generative Models Go MADSina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun 等ICLR 2024 · 被引用 279 次
- Language Modeling Is CompressionGrégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt 等ICLR 2024 · 被引用 243 次
相关 Paper
- Mitigating Source Bias with LLM AlignmentSunhao Dai, Yuqi Zhou, Liang Pang, Zhuoyang Li 等SIGIR 2025 · 被引用 2 次
- Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity DocumentsHaoyu Wang, Sunhao Dai, Haiyuan Zhao, Liang Pang 等ICLR 2025
- Exploring the Escalation of Source Bias in User, Data, and Recommender System Feedback LoopYuqi Zhou, Sunhao Dai, Liang Pang, Gang Wang 等SIGIR 2025 · 被引用 2 次
- AIR-Bench: Automated Heterogeneous Information Retrieval BenchmarkJianlyu Chen, Nan Wang, Chaofan Li, Bo Wang 等ACL 2025
- Attention in Large Language Models Yields Efficient Zero-Shot Re-RankersShijie Chen, Bernal Jimenez Gutierrez, Yu SuICLR 2025
