ValNorm Quantifies Semantics to Reveal Consistent Valence Biases Across Languages and Over Centuries
Autumn Toney, Aylin Caliskan
Abstract
Word embeddings learn implicit biases from linguistic regularities captured by word cooccurrence statistics. By extending methods that quantify human-like biases in word embeddings, we introduce ValNorm, a novel intrinsic evaluation task and method to quantify the valence dimension of affect in human-rated word sets from social psychology. We apply Val-Norm on static word embeddings from seven languages (Chinese, English, German, Polish, Portuguese, Spanish, and Turkish) and from historical English text spanning 200 years. Val-Norm achieves consistently high accuracy in quantifying the valence of non-discriminatory, non-social group word sets. Specifically, Val-Norm achieves a Pearson correlation of ρ = 0.88 for human judgment scores of valence for 399 words collected to establish pleasantness norms in English. In contrast, we measure gender stereotypes using the same set of word embeddings and find that social biases vary across languages. Our results indicate that valence associations of non-discriminatory, non-social group words represent widely-shared associations, in seven languages and over 200 years.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a60ab2e9-313d-439b-aaaf-141a19161c3cCited by top-tier papers5
- Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language ModelsRobert Wolfe, Aylin CaliskanEMNLP 2021 · 27 citations
- VAST: The Valence-Assessing Semantics Test for Contextualizing Language ModelsRobert Wolfe, Aylin CaliskanAAAI 2022 · 17 citations
- Contrastive Visual Semantic Pretraining Magnifies the Semantics of Natural Language RepresentationsRobert Wolfe, Aylin CaliskanACL 2022 · 16 citations
- Reproducibility in Computational Linguistics: Is Source Code Enough?Mohammad Arvan, Luís Pina, Natalie PardeEMNLP 2022 · 12 citations
- The (Undesired) Attenuation of Human Biases by MultilingualityCristina España-Bonet, Alberto Barrón-CedeñoEMNLP 2022 · 5 citations
Related papers
- Gender Bias in Multilingual Embeddings and Cross-Lingual TransferJieyu Zhao, Subhabrata Mukherjee, Saghar Hosseini, Kai-Wei Chang et al.ACL 2020 · 59 citations
- A General Framework for Implicit and Explicit Debiasing of Distributional Word Vector SpacesAnne Lauscher, Goran Glavas, Simone Paolo Ponzetto, Ivan VulicAAAI 2020 · 68 citations
- When do Word Embeddings Accurately Reflect Surveys on our Beliefs About People?Kenneth Joseph, Jonathan H. MorganACL 2020 · 5 citations
- On Measuring and Mitigating Biased Inferences of Word EmbeddingsSunipa Dev, Tao Li, Jeff M. Phillips, Vivek SrikumarAAAI 2020 · 195 citations
- A Causal Inference Method for Reducing Gender Bias in Word Embedding RelationsZekun Yang, Juan FengAAAI 2020 · 40 citations
