The (Undesired) Attenuation of Human Biases by Multilinguality
Cristina España-Bonet, Alberto Barrón-Cedeño
Abstract
Some human preferences are universal. The odor of vanilla is perceived as pleasant all around the world. We expect neural models trained on human texts to exhibit these kind of preferences, i.e. biases, but we show that this is not always the case. We explore 16 static and contextual embedding models in 9 languages and, when possible, compare them under similar training conditions. We introduce and release CA-WEAT, multilingual cultural aware tests to quantify biases, and compare them to previous English-centric tests. Our experiments confirm that monolingual static embeddings do exhibit human biases, but values differ across languages, being far from universal. Biases are less evident in contextual models, to the point that the original human association might be reversed. Multilinguality proves to be another variable that attenuates and even reverses the effect of the bias, specially in contextual multilingual models. In order to explain this variance among models and languages, we examine the effect of asymmetries in the training corpus, departures from isomorphism in multilingual embedding spaces and discrepancies in the testing measures between languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 246bda2b-50f5-4397-80c1-b31ac1b9f8a2Cited by top-tier papers3
- Global Voices, Local Biases: Socio-Cultural Prejudices across LanguagesAnjishnu Mukherjee, Chahat Raj, Ziwei Zhu, Antonios AnastasopoulosEMNLP 2023 · 6 citations
- Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLPPieter Delobelle, Giuseppe Attanasio, Debora Nozza, Su Lin Blodgett et al.EMNLP 2024 · 4 citations
- Semantic and Expressive Variations in Image Captions Across LanguagesAndre Ye, Sebastin Santy, Jena D. Hwang, Amy X. Zhang et al.CVPR 2025
Builds on5
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Null It Out: Guarding Protected Attributes by Iterative Nullspace ProjectionShauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton et al.ACL 2020 · 25 citations
- Are All Good Word Vector Spaces Isomorphic?Ivan Vulic, Sebastian Ruder, Anders SøgaardEMNLP 2020 · 7 citations
- ValNorm Quantifies Semantics to Reveal Consistent Valence Biases Across Languages and Over CenturiesAutumn Toney, Aylin CaliskanEMNLP 2021
- Sense Embeddings are also Biased - Evaluating Social Biases in Static and Contextualised Sense EmbeddingsYi Zhou, Masahiro Kaneko, Danushka BollegalaACL 2022
Related papers
- VAST: The Valence-Assessing Semantics Test for Contextualizing Language ModelsRobert Wolfe, Aylin CaliskanAAAI 2022 · 17 citations
- CARE: Multilingual Human Preference Learning for Cultural AwarenessGeyang Guo, Tarek Naous, Hiromi Wakaki, Yukiko Nishimura et al.EMNLP 2025
- From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association TestXunlian Dai, Li Zhou, Benyou Wang, Haizhou LiEMNLP 2025 · 1 citation
- Gender Bias in Multilingual Embeddings and Cross-Lingual TransferJieyu Zhao, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang et al.ACL 2020 · 59 citations
- SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?Wuttikorn Ponwitayarat, Peerat Limkonchotiwat, Raymond Ng, Jann Railey Montalan et al.ACL 2026 · 2 citations
