Statistical Uncertainty in Word Embeddings: GloVe-V
Andrea Vallebueno, Cassandra Handan-Nader, Christopher D. Manning, Daniel E. Ho
Abstract
Static word embeddings are ubiquitous in computational social science applications and contribute to practical decision-making in a variety of fields including law and healthcare. However, assessing the statistical uncertainty in downstream conclusions drawn from word embedding statistics has remained challenging. When using only point estimates for embeddings, researchers have no streamlined way of assessing the degree to which their model selection criteria or scientific conclusions are subject to noise due to sparsity in the underlying data used to generate the embeddings. We introduce a method to obtain approximate, easy-to-use, and scalable reconstruction error variance estimates for GloVe (Pennington et al., 2014) , one of the most widely used word embedding models, using an analytical approximation to a multivariate normal model. To demonstrate the value of embeddings with variance (GloVe-V), we illustrate how our approach enables principled hypothesis testing in core word embedding tasks, such as comparing the similarity between different word pairs in vector space, assessing the performance of different models, and analyzing the relative degree of ethnic or gender bias in a corpus using different word lists.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on3
- With Little Power Comes Great ResponsibilityDallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia et al.EMNLP 2020 · 76 citations
- RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural PromptsHan Liu, Yuhao Wu, Shixuan Zhai, Bo Yuan et al.CVPR 2023
- Bad Seeds: Evaluating Lexical Methods for Bias MeasurementMaria Antoniak, David MimnoACL 2021
Related papers
- Assessing the Reliability of Word Embedding Gender Bias MeasuresYupei Du, Qixiang Fang, Dong NguyenEMNLP 2021 · 13 citations
- On Measuring and Mitigating Biased Inferences of Word EmbeddingsSunipa Dev, Tao Li, Jeff M. Phillips, Vivek SrikumarAAAI 2020 · 195 citations
- When do Word Embeddings Accurately Reflect Surveys on our Beliefs About People?Kenneth Joseph, Jonathan H. MorganACL 2020 · 5 citations
- ValNorm Quantifies Semantics to Reveal Consistent Valence Biases Across Languages and Over CenturiesAutumn Toney, Aylin CaliskanEMNLP 2021
- VICE: Variational Interpretable Concept EmbeddingsLukas Muttenthaler, Charles Y. Zheng, Patrick McClure, Robert A. Vandermeulen et al.NeurIPS 2022 · 29 citations
