Comparing Biases and the Impact of Multilingual Training across Multiple Languages
Sharon Levy, Neha Anna John, Ling Liu, Yogarshi Vyas, Jie Ma, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, Dan Roth
Abstract
Studies in bias and fairness in natural language processing have primarily examined social biases within a single language and/or across few attributes (e.g. gender, race). However, biases can manifest differently across various languages for individual attributes. As a result, it is critical to examine biases within each language and attribute. Of equal importance is to study how these biases compare across languages and how the biases are affected when training a model on multilingual data versus monolingual data. We present a bias analysis across Italian, Chinese, English, Hebrew, and Spanish on the downstream sentiment analysis task to observe whether specific demographics are viewed more positively. We study bias similarities and differences across these languages and investigate the impact of multilingual vs. monolingual training data. We adapt existing sentiment bias templates in English to Italian, Chinese, Hebrew, and Spanish for four attributes: race, religion, nationality, and gender 1 . Our results reveal similarities in bias expression such as favoritism of groups that are dominant in each language's culture (e.g. majority religions and nationalities). Additionally, we find an increased variation in predictions across protected groups, indicating bias amplification, after multilingual finetuning in comparison to multilingual pretraining. 2 We discuss our choice of binary gender in the Limitations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f28336e8-e9a0-4c12-91af-4e4e9a6c1651Cited by top-tier papers7
- Language Model Tokenizers Introduce Unfairness Between LanguagesAleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, Adel BibiNeurIPS 2023 · 301 citations
- Global Voices, Local Biases: Socio-Cultural Prejudices across LanguagesAnjishnu Mukherjee, Chahat Raj, Ziwei Zhu, Antonios AnastasopoulosEMNLP 2023 · 6 citations
- Delving into Multilingual Ethical Bias: The MSQAD with Statistical Hypothesis Tests for Large Language ModelsSeunguk Yu, Juhwan Choi, YoungBin KimACL 2025 · 2 citations
- Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM CollaborationWeicheng Ma, John J. Guerrerio, Soroush VosoughiEMNLP 2025 · 1 citation
- A Multilingual Social Bias Benchmark Incorporating Thinking ProcessesMasahiro Kaneko, Danushka Bollegala, Timothy BaldwinACL 2026
Builds on10
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than EnglishAurélie Névéol, Yoann Dupont, Julien Bezançon, Karën FortACL 2022 · 61 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
- Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Arbitrary Textual Style Transfer with Small Language ModelsMirac Suzgun, Luke Melas-Kyriazi, Dan JurafskyEMNLP 2022 · 34 citations
- CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsNikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. BowmanEMNLP 2020 · 19 citations
Related papers
- Cross-lingual Transfer Can Worsen Bias in Sentiment AnalysisSeraphina Goldfarb-Tarrant, Björn Ross, Adam LopezEMNLP 2023 · 4 citations
- Gender Bias in Multilingual Embeddings and Cross-Lingual TransferJieyu Zhao, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang et al.ACL 2020 · 59 citations
- Social Bias in Multilingual Language Models: A SurveyLance Calvin Lim Gamboa, Yue Feng, Mark G. LeeEMNLP 2025
- Mitigating Language-Dependent Ethnic Bias in BERTJaimeen Ahn, Alice OhEMNLP 2021 · 5 citations
- Cultural Bias Matters: A Cross-Cultural Benchmark Dataset and Sentiment-Enriched Model for Understanding Multimodal MetaphorsSenqi Yang, Dongyu Zhang, Jing Ren, Ziqi Xu et al.ACL 2025 · 11 citations
