The (Undesired) Attenuation of Human Biases by Multilinguality
Cristina España-Bonet, Alberto Barrón-Cedeño
摘要
Some human preferences are universal. The odor of vanilla is perceived as pleasant all around the world. We expect neural models trained on human texts to exhibit these kind of preferences, i.e. biases, but we show that this is not always the case. We explore 16 static and contextual embedding models in 9 languages and, when possible, compare them under similar training conditions. We introduce and release CA-WEAT, multilingual cultural aware tests to quantify biases, and compare them to previous English-centric tests. Our experiments confirm that monolingual static embeddings do exhibit human biases, but values differ across languages, being far from universal. Biases are less evident in contextual models, to the point that the original human association might be reversed. Multilinguality proves to be another variable that attenuates and even reverses the effect of the bias, specially in contextual multilingual models. In order to explain this variance among models and languages, we examine the effect of asymmetries in the training corpus, departures from isomorphism in multilingual embedding spaces and discrepancies in the testing measures between languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Global Voices, Local Biases: Socio-Cultural Prejudices across LanguagesAnjishnu Mukherjee, Chahat Raj, Ziwei Zhu, Antonios AnastasopoulosEMNLP 2023 · 被引用 6 次
- Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLPPieter Delobelle, Giuseppe Attanasio, Debora Nozza, Su Lin Blodgett 等EMNLP 2024 · 被引用 4 次
- Semantic and Expressive Variations in Image Captions Across LanguagesAndre Ye, Sebastin Santy, Jena D. Hwang, Amy X. Zhang 等CVPR 2025
它引用的顶会 Paper5
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Null It Out: Guarding Protected Attributes by Iterative Nullspace ProjectionShauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton 等ACL 2020 · 被引用 25 次
- Are All Good Word Vector Spaces Isomorphic?Ivan Vulic, Sebastian Ruder, Anders SøgaardEMNLP 2020 · 被引用 7 次
- ValNorm Quantifies Semantics to Reveal Consistent Valence Biases Across Languages and Over CenturiesAutumn Toney, Aylin CaliskanEMNLP 2021
- Sense Embeddings are also Biased - Evaluating Social Biases in Static and Contextualised Sense EmbeddingsYi Zhou, Masahiro Kaneko, Danushka BollegalaACL 2022
相关 Paper
- VAST: The Valence-Assessing Semantics Test for Contextualizing Language ModelsRobert Wolfe, Aylin CaliskanAAAI 2022 · 被引用 17 次
- CARE: Multilingual Human Preference Learning for Cultural AwarenessGeyang Guo, Tarek Naous, Hiromi Wakaki, Yukiko Nishimura 等EMNLP 2025
- From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association TestXunlian Dai, Li Zhou, Benyou Wang, Haizhou LiEMNLP 2025 · 被引用 1 次
- Gender Bias in Multilingual Embeddings and Cross-Lingual TransferJieyu Zhao, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang 等ACL 2020 · 被引用 59 次
- SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?Wuttikorn Ponwitayarat, Peerat Limkonchotiwat, Raymond Ng, Jann Railey Montalan 等ACL 2026 · 被引用 2 次
