Global Voices, Local Biases: Socio-Cultural Prejudices across Languages
Anjishnu Mukherjee, Chahat Raj, Ziwei Zhu, Antonios Anastasopoulos
Abstract
Human biases are ubiquitous but not uniform: disparities exist across linguistic, cultural, and societal borders. As large amounts of recent literature suggest, language models (LMs) trained on human data can reflect and often amplify the effects of these social biases. However, the vast majority of existing studies on bias are heavily skewed towards Western and European languages. In this work, we scale the Word Embedding Association Test (WEAT) to 24 languages, enabling broader studies and yielding interesting findings about LM bias. We additionally enhance this data with culturally relevant information for each language, capturing local contexts on a global scale. Further, to encompass more widely prevalent societal biases, we examine new bias dimensions across toxicity, ableism, and more. Moreover, we delve deeper into the Indian linguistic landscape, conducting a comprehensive regional bias analysis across six prevalent Indian languages. Finally, we highlight the significance of these social biases and the new dimensions through an extensive comparison of embedding methods, reinforcing the need to address them in pursuit of more equitable language models. 1 * Equal contribution 1 All code, data and results are available here: https:// github.com/iamshnoo/weathub .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc17a04a-6160-405e-8e10-26c0f866430eCited by top-tier papers5
- Towards Measuring and Modeling "Culture" in LLMs: A SurveyMuhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh et al.EMNLP 2024 · 21 citations
- The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead?Alexander S. Choi, Syeda Sabrina Akter, JP Singh, Antonios AnastasopoulosEMNLP 2024 · 5 citations
- Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLPPieter Delobelle, Giuseppe Attanasio, Debora Nozza, Su Lin Blodgett et al.EMNLP 2024 · 4 citations
- The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce HarmAakanksha, Arash Ahmadian, Beyza Ermis, Seraphina Goldfarb-Tarrant et al.EMNLP 2024 · 3 citations
- Social Bias in Multilingual Language Models: A SurveyLance Calvin Lim Gamboa, Yue Feng, Mark G. LeeEMNLP 2025
Builds on4
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language ModelsHannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal et al.NeurIPS 2021 · 243 citations
- Comparing Biases and the Impact of Multilingual Training across Multiple LanguagesSharon Levy, Neha Anna John, Ling Liu, Yogarshi Vyas et al.EMNLP 2023 · 9 citations
- The (Undesired) Attenuation of Human Biases by MultilingualityCristina España-Bonet, Alberto Barrón-CedeñoEMNLP 2022 · 5 citations
Related papers
- VAST: The Valence-Assessing Semantics Test for Contextualizing Language ModelsRobert Wolfe, Aylin CaliskanAAAI 2022 · 17 citations
- PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural DataIshaan Watts, Varun Gumma, Aditya Yadavalli, Vivek Seshadri et al.EMNLP 2024 · 3 citations
- Mitigating Language-Dependent Ethnic Bias in BERTJaimeen Ahn, Alice OhEMNLP 2021 · 5 citations
- Digital Skin, Digital Bias: Uncovering Tone-Based Biases in LLMs and Emoji EmbeddingsMingchen Li, Wajdi Aljedaani, Yingjie Liu, Navyasri Meka et al.WWW 2026
- From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association TestXunlian Dai, Li Zhou, Benyou Wang, Haizhou LiEMNLP 2025 · 1 citation
