Bad Seeds: Evaluating Lexical Methods for Bias Measurement
Maria Antoniak, David Mimno
Abstract
Scope of Reproducibility In this work we verify the results of Bad Seeds: Evaluating Lexical Methods for Bias Measurement (Antoniak and Mimno, 2021). We replicate the experiments conducted and verify the main claims made in the original paper: (1) Bias measurements depend on seeds and models. (2) Shuffled seed pairs can result in a significant different bias subspace compared to ordered seed pairs. (3) Set similarity is negatively correlated with the explained variance of the first PCA component in the seed pairings subspace. Methodology We used skip-gram with negative sampling to train word2vec models with the same hyperparameters and data. We implemented code for the experiments using the resulting word embeddings and the seed sets provided by the authors. Results Overall, only one claim was reproduced. We reproduced the claim that bias measurements is dependent on the choice of seed set. We were not able to adequately reproduce the claims that shuffled pairs of seed sets generally result in less clearly defined correlation and that for pair of seed sets set similarity is negatively correlated with the explained variance of the first principal component. What was easy The paper is easy to follow. The data was publicly available. Authors replied frequently providing details about the parameters and preprocessing steps. Also authors were open for the discussion regarding seeds on the Github repository of the project. What was difficult In certain cases, the gathered seed sets json file contained errors. Specifically: 'daughters' was misspelled as 'daughers' (has now been updated). 'ma', 'am' was used as two words instead of one word "ma'am". The seed words for figure 4 as stated in the appendix are not the same as the one in the image itself. Table 2 from the original paper was also difficult to reproduce as the preprocessing according to the authors' description gave close but not equal results for the NYT dataset and significantly different results for two other datasets. Reproduced numbers are presented in table 3.3. Because of the time constraints, 20 bootstrapped launches were not conducted. Communication with original authors Contact was made with the original authors on multiple occasions to ask for clarification questions regarding the implementation of the experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language ModelsHannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal et al.NeurIPS 2021 · 243 citations
- Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?Rishi Bommasani, Kathleen A. Creel, Ananya Kumar, Dan Jurafsky et al.NeurIPS 2022 · 179 citations
- Faithful Explanations of Black-box NLP Models Using LLM-generated CounterfactualsYair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin et al.ICLR 2024 · 55 citations
- Rethinking "Risk" in Algorithmic Systems Through A Computational Narrative Analysis of Casenotes in Child-WelfareDevansh Saxena, Erina Seh-Young Moon, Aryan Chaurasia, Yixin Guan et al.CHI 2023 · 35 citations
- Under the Morphosyntactic Lens: A Multifaceted Evaluation of Gender Bias in Speech TranslationBeatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri et al.ACL 2022 · 30 citations
Related papers
- Assessing the Reliability of Word Embedding Gender Bias MeasuresYupei Du, Qixiang Fang, Dong NguyenEMNLP 2021 · 13 citations
- Deep Fair Clustering for Visual LearningPeizhao Li, Han Zhao, Hongfu LiuCVPR 2020
- Intrinsic Bias Metrics Do Not Correlate with Application BiasSeraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sánchez, Mugdha Pandya et al.ACL 2021
- Reproducibility Issues for BERT-based Evaluation MetricsYanran Chen, Jonas Belouadi, Steffen EgerEMNLP 2022 · 11 citations
- Towards Robustifying NLI Models Against Lexical Dataset BiasesXiang Zhou, Mohit BansalACL 2020 · 36 citations
