Media Source Matters More Than Content: Unveiling Political Bias in LLM-Generated Citations
Sunhao Dai, Zhanshuo Cao, Wenjie Wang, Liang Pang, Jun Xu, See-Kiong Ng, Tat-Seng Chua
摘要
Unlike traditional search engines that present ranked lists of webpages, generative search engines rely solely on in-line citations as the key gateway to original real-world webpages, making it crucial to examine whether LLMgenerated citations have biases-particularly for politically sensitive queries. To investigate this, we first construct AllSides-2024, a new dataset comprising the latest real-world news articles (Jan. 2024 -Dec. 2024) labeled with left-or right-leaning stances. Through systematic evaluations, we find that LLMs exhibit a consistent tendency to cite left-leaning sources at notably higher rates compared to traditional retrieval systems (e.g., BM25 and dense retrievers). Controlled experiments further reveal that this bias arises from a preference for media outlets identified as left-leaning, rather than for left-oriented content itself. Meanwhile, our findings show that while LLMs struggle to infer political bias from news content alone, they can almost perfectly recognize the political orientation of media outlets based on their names. These insights highlight the risk that, in the era of generative search engines, information exposure may be disproportionately shaped by specific media outlets, potentially shaping public perception and decision-making 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Whose Facts Win? LLM Source Preferences under Knowledge ConflictsJakob Schuster, Vagrant Gautam, Katja MarkertACL 2026 · 被引用 3 次
- How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI OverviewsRiley Grossman, Songjiang Liu, Michael K. Chen, Mike Smith 等SIGIR 2026
它引用的顶会 Paper9
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin 等SIGIR 2021 · 被引用 297 次
- From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsShangbin Feng, Chan Young Park, Yuhan Liu, Yulia TsvetkovACL 2023 · 被引用 117 次
- Fake News in Sheep's Clothing: Robust Fake News Detection Against LLM-Empowered Style AttacksJiaying Wu, Jiafeng Guo, Bryan HooiKDD 2024 · 被引用 69 次
- RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-EncoderShitao Xiao, Zheng Liu, Yingxia Shao, Zhao CaoEMNLP 2022 · 被引用 63 次
- This Is Not What We Ordered: Exploring Why Biased Search Result Rankings Affect User Attitudes on Debated TopicsTim Draws, Nava Tintarev, Ujwal Gadiraju, Alessandro Bozzon 等SIGIR 2021 · 被引用 36 次
相关 Paper
- Assessing Reliability and Political Bias In LLMs' Judgements of Formal and Material Inferences With Partisan ConclusionsReto Gubelmann, Ghassen KarrayACL 2025
- Fair or Framed? Political Bias in News Articles Generated by LLMsJunho Yoo, Youhyun ShinEMNLP 2025 · 被引用 1 次
- Measuring and Mitigating Media Outlet Name Bias in Large Language ModelsSeong-Jin Park, Kang-Min KimEMNLP 2025
- In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM GenerationsMohammad Aflah Khan, Mahsa Amani, Soumi Das, Bishwamittra Ghosh 等ICLR 2026 · 被引用 4 次
- Generative Echo Chamber? Effect of LLM-Powered Search Systems on Diverse Information SeekingNikhil Sharma, Q. Vera Liao, Ziang XiaoCHI 2024 · 被引用 123 次
