When Bigger Isn't Better: A Comprehensive Fairness Evaluation of Political Bias in Multi-News Summarisation
Nannan Huang, Iffat Maab, Junichi Yamagishi
Abstract
Multi-document news summarisation systems are increasingly adopted for their convenience in processing vast daily news content, making fairness across diverse political perspectives critical. However, these systems can exhibit political bias through unequal representation of viewpoints, disproportionate emphasis on certain perspectives, and systematic underrepresentation of minority voices. This study presents a comprehensive evaluation of such bias in multi-document news summarisation using FairNews, a dataset of complete news articles with political orientation labels, examining how large language models (LLMs) handle sources with varying political leanings across 13 models and five fairness metrics. We investigate both baseline model performance and effectiveness of various debiasing interventions, including prompt-based and judge-based approaches. Our findings challenge the assumption that larger models yield fairer outputs, as mid-sized variants consistently outperform their larger counterparts, offering the best balance of fairness and efficiency. Prompt-based debiasing proves highly model dependent, while entity sentiment emerges as the most stubborn fairness dimension, resisting all intervention strategies tested. These results demonstrate that fairness in multi-document news summarisation requires multi-dimensional evaluation frameworks and targeted, architecture-aware debiasing rather than simply scaling up.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed9b3219-5c4e-4e41-8b8b-24522480844cBuilds on7
- Co-Writing with Opinionated Language Models Affects Users' ViewsMaurice Jakesch, Advait Bhat, Daniel Buschek, Lior Zalmanson et al.CHI 2023 · 249 citations
- From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsShangbin Feng, Chan Young Park, Yuhan Liu, Yulia TsvetkovACL 2023 · 117 citations
- "Thinking" Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language ModelsShaz Furniturewala, Surgan Jandial, Abhinav Java, Pragyan Banerjee et al.EMNLP 2024 · 14 citations
- We Can Detect Your Bias: Predicting the Political Ideology of News ArticlesRamy Baly, Giovanni Da San Martino, James R. Glass, Preslav NakovEMNLP 2020 · 6 citations
- Less Is More? Examining Fairness in Pruned Large Language Models for Summarising OpinionsNannan Huang, Haytham M. Fayek, Xiuzhen ZhangEMNLP 2025 · 2 citations
Related papers
- Fair or Framed? Political Bias in News Articles Generated by LLMsJunho Yoo, Youhyun ShinEMNLP 2025 · 1 citation
- Measuring and Mitigating Media Outlet Name Bias in Large Language ModelsSeong-Jin Park, Kang-Min KimEMNLP 2025
- One Prompt To Rule Them All: LLMs for Opinion Summary EvaluationTejpalsingh Siledar, Swaroop Nath, Sankara Sri Raghava Ravindra Muddu, Rupasai Rangaraju et al.ACL 2024
- FAIR-RAG: An End-to-End Framework for Mitigating Political Bias through Fair Retrieval-Augmented GenerationJaebeom You, Kisung Lee, Hyuk-Yoon KwonSIGIR 2026
- LLMS ON TRIAL: Evaluating Judicial Fairness For Large Language ModelsYiran Hu, Zongyue Xue, Haitao Li, Siyuan Zheng et al.ICLR 2026 · 4 citations
