BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation
Tianxiang Sun, Junliang He, Xipeng Qiu, Xuanjing Huang
Abstract
WARNING: This paper contains examples that are offensive in nature. Automatic evaluation metrics are crucial to the development of generative systems. In recent years, pre-trained language model (PLM) based metrics, such as BERTScore (Zhang et al., 2020) , have been commonly adopted in various generation tasks. However, it has been demonstrated that PLMs encode a range of stereotypical societal biases, leading to a concern on the fairness of PLMs as metrics. To that end, this work presents the first systematic study on the social bias in PLM-based metrics. We demonstrate that popular PLMbased metrics exhibit significantly higher social bias than traditional metrics on 6 sensitive attributes, namely race, gender, religion, physical appearance, age, and socioeconomic status. In-depth analysis suggests that choosing paradigms (matching, regression, or generation) of the metric has a greater impact on fairness than choosing PLMs. In addition, we develop debiasing adapters that are injected into PLM layers, mitigating bias in PLM-based metrics while retaining high performance for evaluating text generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d9092be-5308-478d-8993-edfe4f8844d3Cited by top-tier papers6
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- Measuring Political Bias in Large Language Models: What Is Said and How It Is SaidYejin Bang, Delong Chen, Nayeon Lee, Pascale FungACL 2024 · 21 citations
- Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality EstimationEmmanouil Zaranis, Giuseppe Attanasio, Sweta Agrawal, André F. T. MartinsACL 2025 · 8 citations
- Framing Political Bias in Multilingual LLMs Across Pakistani LanguagesAfrozah Nadeem, Mark Dras, Usman NaseemACL 2026 · 6 citations
- Gender Biases in Automatic Evaluation Metrics for Image CaptioningHaoyi Qiu, Zi-Yi Dou, Tianlu Wang, Asli Celikyilmaz et al.EMNLP 2023 · 6 citations
Builds on15
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
Related papers
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 495 citations
- Debiasing Pretrained Text Encoders by Paying Attention to Paying AttentionYacine Gaci, Boualem Benatallah, Fabio Casati, Khalid BenabdeslemEMNLP 2022 · 12 citations
- Auto-Debias: Debiasing Masked Language Models with Automated Biased PromptsYue Guo, Yi Yang, Ahmed AbbasiACL 2022
- RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language ModelsSoumya Barikeri, Anne Lauscher, Ivan Vulic, Goran GlavasACL 2021
- Mitigating Language-Dependent Ethnic Bias in BERTJaimeen Ahn, Alice OhEMNLP 2021 · 5 citations
