Mitigating Source Bias with LLM Alignment
Sunhao Dai, Yuqi Zhou, Liang Pang, Zhuoyang Li, Zhaocheng Du, Gang Wang, Jun Xu
Abstract
Recent studies have revealed a phenomenon known as source bias, where PLM-based retrievers assign higher relevance scores to LLM-generated content despite its semantic quality being comparable to human-written content. As LLMs rapidly advance and become more widely used, effectively counteracting source bias is crucial for the sustainable development of the information retrieval (IR) ecosystem. Existing methods primarily attempt to address source bias from the retriever side, adopting a "passive defense" approach that intervenes only after biased content has entered the retrieval pipeline. These solutions are limited by frequent retriever updates in industrial applications, high recurring costs, and their inability to address the root cause of source bias.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e91365d1-b8a0-4db7-9891-d8b0e327e943Cited by top-tier papers2
- Exploring the Escalation of Source Bias in User, Data, and Recommender System Feedback LoopYuqi Zhou, Sunhao Dai, Liang Pang, Gang Wang et al.SIGIR 2025 · 2 citations
- How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI OverviewsRiley Grossman, Songjiang Liu, Michael K. Chen, Mike Smith et al.SIGIR 2026
Related papers
- Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity DocumentsHaoyu Wang, Sunhao Dai, Haiyuan Zhao, Liang Pang et al.ICLR 2025
- Neural Retrievers are Biased Towards LLM-Generated ContentSunhao Dai, Yuqi Zhou, Liang Pang, Weihao Liu et al.KDD 2024 · 26 citations
- Invisible Relevance Bias: Text-Image Retrieval Models Prefer AI-Generated ImagesShicheng Xu, Danyang Hou, Liang Pang, Jingcheng Deng et al.SIGIR 2024 · 18 citations
- LLM Alignment as Retriever Optimization: An Information Retrieval PerspectiveBowen Jin, Jinsung Yoon, Zhen Qin, Ziqi Wang et al.ICML 2025
- Blinded by Generated Contexts: How Language Models Merge Generated and Retrieved Contexts When Knowledge Conflicts?Hexiang Tan, Fei Sun, Wanli Yang, Yuanzhuo Wang et al.ACL 2024
